# #grpo — MDRSS hashtag feed

> Public MDRSS cards tagged #grpo.
> Canonical feed: https://mdrss.com/feeds/grpo

## Cards (3)

### [GRPO Reward Function Library](https://mdrss.com/llm-engineering/training-and-fine-tuning/901369/901369.md)

Complete, runnable reward functions for TRL's GRPOTrainer. Every function here follows the current TRL reward-function signature: it accepts completions plus any extra dataset columns as keyword arguments, and returns a list float  the same length as completions. Base mod. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: llm-engineering/training-and-fine-tuning · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [GRPO & RLVR Training](https://mdrss.com/llm-engineering/training-and-fine-tuning/901368/901368.md)

This skill assumes finetuning-method-selection already routed here because the target behavior has a verifiable pass/fail signal — not demonstrations (lora-qlora-recipes) or preference pairs (preference-optimization). What follows is when RL is the right tool, the reference. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: llm-engineering/training-and-fine-tuning · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale](https://mdrss.com/ai-agents/agent-frameworks/1549/1549.md)

Towards Async, Omni-Modal RL at Scale, Just Relax. 📖 English | 📖 中文 Relax (Reinforcement Engine Leveraging Agentic X-modality) is a high-performance reinforcement learning post-training framework open-sourced by the Xiaohongshu AI Infra Team for multimodal large language models.

Classification: ai-agents/agent-frameworks · Feed: ai-agents · Updated: 2026-08-04T12:22:38.168Z · Version: 1
