Explore indexBack to Terms
Direct Preference Optimization
Direct Preference Optimization (DPO) aligns language models directly on pairwise human preference data (chosen vs. rejected responses) without requiring an explicit reward model or complex reinforcement learning loops. It stabilizes policy training for refining tone, conciseness, safety refusal thresholds, and tool selection preferences.
No public content is connected to this entity yet.