CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of Direct Preference Optimization (DPO) in large language model (LLM) planning, specifically its neglect of constraint violation severity and inherent biases in preference data. To overcome these issues, we propose CM-DPO, a method that leverages a symbolic verifier to generate continuous constraint margin signals and employs a lexicographic objective function to strictly decouple hard and soft constraints. Furthermore, we introduce the SynPlan-R framework, which integrates programmatic generation with minimal-edit distillation to construct unbiased training data. Experimental results demonstrate that an 8B-parameter model achieves an 89.2% pass rate while reducing inference latency by 13 times. Notably, our approach outperforms GPT-4o by 9.2 percentage points on the Blocksworld benchmark, highlighting its effectiveness for efficient and reliable LLM-based planning.
📝 Abstract
Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.
Problem

Research questions and friction points this paper is trying to address.

Direct Preference Optimization
Constraint Violation
LLM Planning
Length and Style Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constraint-Margin DPO
Lexicographic Objective
Symbolic Verifier
Preference Data Synthesis
LLM Planning
💼 Related Jobs
No related jobs found.
R
Rabimba Karanjai
PayPal AI Lab
Q
Qun Gu
PayPal AI Lab
H
Hemanth Hegadehalli Madhavarao
PayPal AI Lab
Wenhuan Sun
Wenhuan Sun
PayPal AI Lab
Xiaojiao Yu
Xiaojiao Yu
PayPal AI Lab
Suryabhan Singh Hada
Suryabhan Singh Hada
PayPal AI Lab
L
Libin N. George
PayPal AI Lab
U
Uma Kona
PayPal AI Lab
R
Richard Williamson
PayPal AI Lab
L
Linsey Pang
PayPal AI Lab
P
Prakhar Mehrotra
PayPal AI Lab