Intermediate1.5 to 2 hours
Preference Optimization After DPO: IPO, KTO, and SimPO Compared
Runnable comparison of preference optimization after DPO: IPO, KTO, and SimPO. What each objective changes, when reference-free wins, and how to switch in TRL.
Prerequisites
Comfort with the Hugging Face stack (transformerstrlpeftdatasets)one 24GB+ GPUand prior exposure to supervised fine-tuning. Understanding of DPO at a conceptual level helps but is not required.