🤖 AI Summary
This study addresses the issue in multi-turn dialogues where rejected or replaced historical intents continue to interfere with large language model decision-making, leading to task failure. To tackle the model's difficulty in distinguishing between "mentioned" and "active" intents, this work proposes the Intent-Eval benchmark and the Intent-OPSD framework. Built upon online policy self-distillation, the framework initializes teacher and student networks from a homogeneous model and employs decision-conditioned supervision with a frozen teacher to precisely guide the model toward focusing on currently active intents. Experimental results demonstrate that the proposed method significantly enhances robustness in scenarios such as tool calling and code generation, effectively mitigating performance degradation caused by irrelevant historical dialogue turns.
📝 Abstract
When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics. Across diverse tasks, models are vulnerable to both rejected proposals and superseded requirements, consistent with mentioned-as-in-effect confusion: conversational content is treated as active requirements even after it has been rejected or replaced. Accuracy degradation can deepen or persist as interaction continues, highlighting the need to distinguish what has been mentioned from what remains in effect. Building on this insight, we propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework with Teacher and Student initialized from the same model. The frozen Teacher provides active-intent supervision from the complete task matching the user's decision, training the Student on the full dialogue to follow active requirements reflecting user intent.