MDRSS · MARKDOWN SNAPSHOT
9.8/10

GRPO & RLVR Training

This skill assumes finetuning-method-selection already routed here because the target behavior has a verifiable pass/fail signal — not demonstrations (lora-qlora-recipes) or preference pairs (preference-optimization). What follows is when RL is the right tool, the reference. Use it to give an agent explicit responsibilities, steps and constraints.

#llm-engineering / card #901368★ 0◌ 0snapshot 2026-08-04