Bidirectional Preference Synthesis: Learning Prompt-Conditioned Preferences from Boundary Failures

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of supervision for boundary failure samples in offline preference optimization, where responses violate the original instruction yet satisfy semantically adjacent intents. To this end, we propose a bidirectional preference synthesis method that constructs forward and reverse paired data, introducing reverse preference pairs under "achieved prompts" so that identical responses are rejected in incorrect contexts while selected in correct ones. This explicitly models prompt-conditioned dependencies, refining supervision signals without modifying the DPO objective or training a reward model. Experiments demonstrate that our approach improves achieved-side ranking accuracy from 6.8% to 62.3%, significantly enhancing multilingual multi-turn instruction-following capabilities while maintaining consistent advantages across agent-based, tool-calling, and code generation tasks.
📝 Abstract
Correction-based offline preference pipelines commonly treat model failures only as rejected responses under the original prompt. This supervision is incomplete for boundary failures: responses that violate the given instruction yet coherently satisfy a nearby intent or constraint setting. We introduce Bidirectional Preference Synthesis (BPS), a data-construction method for standard Direct Preference Optimization (DPO) that makes this missing prompt dependence explicit. For each validated boundary failure, BPS keeps the conventional forward pair under the original prompt and adds a reverse pair under a synthesized achieved prompt, so the same response is rejected where it is wrong and chosen where it is right, without changing the DPO objective, training a reward model, or requiring online sampling. On Qwen3-4B-Instruct-2507, BPS preserves original-side pairwise ranking while raising achieved-side ranking accuracy from 6.8% to 62.3% on held-out crossed anchors, with a similar shift under a Kimi-K2.6 cross-teacher probe. A blind human audit supports the intended reverse preference direction, and downstream evaluations show the clearest separation from Forward-DPO in multilingual multi-turn instruction following, with consistent capability-retention patterns on agentic, tool-use, and code checks.
Problem

Research questions and friction points this paper is trying to address.

offline preference learning
boundary failures
prompt-conditioned preferences
Direct Preference Optimization
instruction following
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bidirectional Preference Synthesis
Boundary Failures
Direct Preference Optimization
Prompt-Conditioned Preferences
Offline Preference Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junbo Wang
Kuaishou Technology
Lidong Lu
Lidong Lu
Nanjing University
Multimodal Large Language Model
Zhuoqun Li
Zhuoqun Li
Institute of Software, Chinese Academy of Sciences
Natural Language Processing
G
Guiping Jiang
Kuaishou Technology
X
Xiangyu Wu
Kuaishou Technology
T
Tinghai Zhang
Kuaishou Technology
Tong Lu
Tong Lu
Nanjing University
Computer VisionFoundation Models