When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure mode in which LLM-based agents, during multi-turn interactions, suffer degraded output quality due to historical instructions interfering with evolving user intents. We formally define and quantify this phenomenon as "intent drift," and construct IntentFlux, an executable benchmark designed to evaluate it. To mitigate this issue, we propose StateForge, a method that explicitly tracks and maintains active requirement states to suppress interference from obsolete instructions. Experimental results demonstrate that intent drift significantly degrades task performance, while StateForge effectively alleviates this degradation by improving the average score from 0.367 to 0.467. This work provides a principled state-maintenance mechanism for robust multi-turn dialogue systems.
📝 Abstract
LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders. In a 627-case calibration, mean task score falls from 0.476 to 0.384 as dialogues contain more superseded and withdrawn information. Across eight models, the rate of fully correct solutions is significantly lower when the same final task must be recovered from an evolving dialogue rather than given directly in a single turn. We further introduce StateForge, which explicitly maintains the active requirements before generation. On General-Test, it improves mean task score from 0.367 to 0.467. Providing the ground-truth final state improves performance further but still does not recover single-turn performance, indicating that state-estimation errors explain only part of the gap. These results establish intent drift as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation.
Problem

Research questions and friction points this paper is trying to address.

Intent Drift
LLM Agents
Multi-turn Interaction
User Intent Change
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intent Drift
IntentFlux Benchmark
StateForge
LLM Agents
Explicit State Maintenance
🔎 Similar Papers