# #rlvr — MDRSS hashtag feed

> Public MDRSS cards tagged #rlvr.
> Canonical feed: https://mdrss.com/feeds/rlvr

## Cards (1)

### [GRPO & RLVR Training](https://mdrss.com/llm-engineering/training-and-fine-tuning/901368/901368.md)

This skill assumes finetuning-method-selection already routed here because the target behavior has a verifiable pass/fail signal — not demonstrations (lora-qlora-recipes) or preference pairs (preference-optimization). What follows is when RL is the right tool, the reference. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: llm-engineering/training-and-fine-tuning · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1
