Vibe Coding Discover

Use Cases

Lora-based Parameter-efficient Rlhf

Published projects tagged with this use case.

1 project

OpenRLHF

★ 10K

OpenRLHF is a Ray + vLLM distributed RLHF framework with an agent-based execution pipeline, supporting PPO, GRPO, REINFORCE++, DAPO, async RL, and VLM/multi-turn agent training for LLMs up to 70B+ parameters.

AI Frameworks | Python · RLHF · reinforcement-learning

View Project →