openrlhf-training
High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/openrlhf-training/SKILL.md
- skills/openrlhf-training/references/algorithm-comparison.md
- skills/openrlhf-training/references/custom-rewards.md
- skills/openrlhf-training/references/hybrid-engine.md
- skills/openrlhf-training/references/multi-node-training.md
Every link opens the file at its source, pinned to the revision this page describes.