MDRSS ยท MARKDOWN SNAPSHOT
9.8/10

Preference Optimization

This skill assumes finetuning-method-selection already routed here because the data shape is preference pairs or unpaired thumbs-up/down feedback, not demonstrations (that's lora-qlora-recipes) or a verifiable reward signal (that's grpo-rlvr-training). What follows is metho. Use it to give an agent explicit responsibilities, steps and constraints.

#llm-engineering / card #901373โ˜… 0โ—Œ 0snapshot 2026-08-04