fine-tuning-with-trl
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/fine-tuning-with-trl/SKILL.md
- skills/fine-tuning-with-trl/references/dpo-variants.md
- skills/fine-tuning-with-trl/references/online-rl.md
- skills/fine-tuning-with-trl/references/reward-modeling.md
- skills/fine-tuning-with-trl/references/sft-training.md
Every link opens the file at its source, pinned to the revision this page describes.