🤖 AI Summary
This study addresses the lack of supervision for boundary failure samples in offline preference optimization, where responses violate the original instruction yet satisfy semantically adjacent intents. To this end, we propose a bidirectional preference synthesis method that constructs forward and reverse paired data, introducing reverse preference pairs under "achieved prompts" so that identical responses are rejected in incorrect contexts while selected in correct ones. This explicitly models prompt-conditioned dependencies, refining supervision signals without modifying the DPO objective or training a reward model. Experiments demonstrate that our approach improves achieved-side ranking accuracy from 6.8% to 62.3%, significantly enhancing multilingual multi-turn instruction-following capabilities while maintaining consistent advantages across agent-based, tool-calling, and code generation tasks.
📝 Abstract
Correction-based offline preference pipelines commonly treat model failures only as rejected responses under the original prompt. This supervision is incomplete for boundary failures: responses that violate the given instruction yet coherently satisfy a nearby intent or constraint setting. We introduce Bidirectional Preference Synthesis (BPS), a data-construction method for standard Direct Preference Optimization (DPO) that makes this missing prompt dependence explicit. For each validated boundary failure, BPS keeps the conventional forward pair under the original prompt and adds a reverse pair under a synthesized achieved prompt, so the same response is rejected where it is wrong and chosen where it is right, without changing the DPO objective, training a reward model, or requiring online sampling. On Qwen3-4B-Instruct-2507, BPS preserves original-side pairwise ranking while raising achieved-side ranking accuracy from 6.8% to 62.3% on held-out crossed anchors, with a similar shift under a Kimi-K2.6 cross-teacher probe. A blind human audit supports the intended reverse preference direction, and downstream evaluations show the clearest separation from Forward-DPO in multilingual multi-turn instruction following, with consistent capability-retention patterns on agentic, tool-use, and code checks.