🤖 AI Summary
This study addresses the limitations of Direct Preference Optimization (DPO) in large language model (LLM) planning, specifically its neglect of constraint violation severity and inherent biases in preference data. To overcome these issues, we propose CM-DPO, a method that leverages a symbolic verifier to generate continuous constraint margin signals and employs a lexicographic objective function to strictly decouple hard and soft constraints. Furthermore, we introduce the SynPlan-R framework, which integrates programmatic generation with minimal-edit distillation to construct unbiased training data. Experimental results demonstrate that an 8B-parameter model achieves an 89.2% pass rate while reducing inference latency by 13 times. Notably, our approach outperforms GPT-4o by 9.2 percentage points on the Blocksworld benchmark, highlighting its effectiveness for efficient and reliable LLM-based planning.
📝 Abstract
Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.