OpenRLHF
★ 10KOpenRLHF is a Ray + vLLM distributed RLHF framework with an agent-based execution pipeline, supporting PPO, GRPO, REINFORCE++, DAPO, async RL, and VLM/multi-turn agent training for LLMs up to 70B+ parameters.
AI Frameworks | Python · RLHF · reinforcement-learning
View Project →