MDRSS ยท MARKDOWN SNAPSHOT
9.8/10

Vision-Language SFT

This skill assumes finetuning-method-selection already routed here: the data shape is image+text demonstrations, not preference pairs or a verifiable reward signal, and the base is a vision-language model rather than a text-only one. lora-qlora-recipes covers the text-only Lo. Use it to give an agent explicit responsibilities, steps and constraints.

#llm-engineering / card #901378โ˜… 0โ—Œ 0snapshot 2026-08-04