Online-RLHF
★ 546Recipe and scripts for online iterative RLHF: SFT, reward modeling, vLLM response generation, reward annotation, and iterative DPO training loops. Reproduces LLaMA3-8B alignment comparable to Llama3-8B-Instruct using only open-source data.
AI Frameworks | Python · rlhf · dpo
View Project →