🤖 AI Summary
This work addresses a critical limitation in existing large language model systems, where control decisions—such as whether to answer, seek clarification, or invoke external tools—are tightly coupled with text generation, hindering failure diagnosis. To resolve this, the authors propose a decision-centric framework that explicitly models control decisions as a standalone module, decoupling evaluation from action through a signal–policy separation architecture. This design enables precise failure attribution and modular system improvement. The framework unifies handling of both single-step and sequential decisions and integrates mechanisms such as routing and adaptive reasoning to achieve interpretable and intervenable control. Experimental results demonstrate that the approach significantly reduces ineffective actions and improves task success rates across three benchmarks, while also revealing distinct and diagnosable failure patterns.
📝 Abstract
LLM systems must make control decisions in addition to generating outputs: whether to answer, clarify, retrieve, call tools, repair, or escalate. In many current architectures, these decisions remain implicit within generation, entangling assessment and action in a single model call and making failures hard to inspect, constrain, or repair. We propose a decision-centric framework that separates decision-relevant signals from the policy that maps them to actions, turning control into an explicit and inspectable layer of the system. This separation supports attribution of failures to signal estimation, decision policy, or execution, and enables modular improvement of each component. It unifies familiar single-step settings such as routing and adaptive inference, and extends naturally to sequential settings in which actions alter the information available before acting. Across three controlled experiments, the framework reduces futile actions, improves task success, and reveals interpretable failure modes. More broadly, it offers a general architectural principle for building more reliable, controllable, and diagnosable LLM systems.