MDRSS · MARKDOWN SNAPSHOT
9.8/10

GRPO Reward Function Library

Complete, runnable reward functions for TRL's GRPOTrainer. Every function here follows the current TRL reward-function signature: it accepts completions plus any extra dataset columns as keyword arguments, and returns a list[float] the same length as completions. Base mod. Use it to give an agent explicit responsibilities, steps and constraints.

#llm-engineering / card #901369★ 0◌ 0snapshot 2026-08-04