# #rlhf — MDRSS hashtag feed

> Public MDRSS cards tagged #rlhf.
> Canonical feed: https://mdrss.com/feeds/rlhf

## Cards (2)

### [DLLM RL](https://mdrss.com/llm-engineering/serving-and-retrieval/1891/1891.md)

We also introduce a diffusion-based value model that reduces variance and improves stability during optimization. Based on TraceRL, we derive a series of diffusion language models, TraDo, which achieve state-of-the-art performance on math and coding reasoning tasks.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [MLX-LM-LORA](https://mdrss.com/llm-engineering/models-and-training/1483/1483.md)

With MLX-LM-LoRA you can, train Large Language Models locally on Apple Silicon using MLX. Training works with all models supported by MLX-LM, including: Training Types: Training Algorithms: Quantization Aware Training (QAT): Training Your Custom Preference Model: --- The main command is mlxlmlora.train.

Classification: llm-engineering/models-and-training · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1
