{"version":"mdrss-hashtag-feed/1","tag":"grpo","urls":{"html":"https://mdrss.com/feeds/grpo","rss":"https://mdrss.com/feeds/grpo/rss.xml","json":"https://mdrss.com/feeds/grpo/feed.json","markdown":"https://mdrss.com/feeds/grpo/index.md"},"updated_at":"2026-08-04T13:54:51.641Z","items":[{"schema":"mdrss.card-summary/v1","id":901369,"version":1,"title":"GRPO Reward Function Library","annotation":"Complete, runnable reward functions for TRL's GRPOTrainer. Every function here follows the current TRL reward-function signature: it accepts completions plus any extra dataset columns as keyword arguments, and returns a list[float] the same length as completions. Base mod. Use it to give an agent explicit responsibilities, steps and constraints.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"training-and-fine-tuning","content_type":"guide","tags":["llm-ml-engineering","training-fine-tuning","reward","function","correctness","grpo","library","training","llm-engineering","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:45.280Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901369","permalink_url":"https://mdrss.com/m/901369","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901369/901369.md","file_url":"https://mdrss.com/api/v1/cards/901369/file","raw_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901369/raw","embed_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901369/embed","edit_url":"https://mdrss.com/cards/901369/edit","legacy_url":"https://mdrss.com/s/llm-engineering/grpo-reward-function-library-collider-a43d7fc781e8"}},{"schema":"mdrss.card-summary/v1","id":901368,"version":1,"title":"GRPO & RLVR Training","annotation":"This skill assumes finetuning-method-selection already routed here because the target behavior has a verifiable pass/fail signal — not demonstrations (lora-qlora-recipes) or preference pairs (preference-optimization). What follows is when RL is the right tool, the reference. Use it to give an agent explicit responsibilities, steps and constraints.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"training-and-fine-tuning","content_type":"guide","tags":["llm-ml-engineering","training-fine-tuning","grpo","rlvr","training","fine-tuning","llm","ml","llm-engineering","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:45.280Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901368","permalink_url":"https://mdrss.com/m/901368","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901368/901368.md","file_url":"https://mdrss.com/api/v1/cards/901368/file","raw_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901368/raw","embed_url":"https://mdrss.com/llm-engineering/training-and-fine-tuning/901368/embed","edit_url":"https://mdrss.com/cards/901368/edit","legacy_url":"https://mdrss.com/s/llm-engineering/grpo-rlvr-training-collider-204e909c2dea"}},{"schema":"mdrss.card-summary/v1","id":1549,"version":1,"title":"Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale","annotation":"Towards Async, Omni-Modal RL at Scale, Just Relax. 📖 English | 📖 中文 Relax (Reinforcement Engine Leveraging Agentic X-modality) is a high-performance reinforcement learning post-training framework open-sourced by the Xiaohongshu AI Infra Team for multimodal large language models.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"agent-frameworks","content_type":"guide","tags":["agentic-rl","distributed-training","grpo","llm","megatron-lm","multi-agent","multimodal","post-training"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:15:38.483Z","urls":{"card_url":"https://mdrss.com/ai-agents/agent-frameworks/1549","permalink_url":"https://mdrss.com/m/1549","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/agent-frameworks/1549/1549.md","file_url":"https://mdrss.com/api/v1/cards/1549/file","raw_url":"https://mdrss.com/ai-agents/agent-frameworks/1549/raw","embed_url":"https://mdrss.com/ai-agents/agent-frameworks/1549/embed","edit_url":"https://mdrss.com/cards/1549/edit","legacy_url":"https://mdrss.com/s/ai-agents/redai-infra-relax-redai-infra-relax-readme"}}]}