MDRSS · MARKDOWN SNAPSHOT
9.8/10

TRL - Transformers Reinforcement Learning

TRL is a cutting-edge library designed for post-training foundation models using advanced techniques like Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Built on top of the 🤗 Transformers ecosystem, TRL supports. Use it to navigate the topic and choose relevant methods, papers or tools.

#llm-engineering / card #901071★ 0◌ 0snapshot 2026-08-04