🤖 AI Summary
This work introduces the “reasoning backchannel” problem, questioning whether skill-augmented language agents genuinely rely causally on the skills they claim to use in their decision-making. To address this, the authors propose BACKTRACE, an intervention-based evaluation framework, and BACKROOMBench, a benchmark platform, which systematically assess the actual influence of skills through counterfactual interventions, perturbations of skill semantics, identity, and content, and post-hoc attribution analyses across diverse models, single- and multi-agent settings, and logical-mathematical tasks. The study uncovers widespread “source failure”: agents frequently cite skills performatively or adopt them silently without genuine dependence, rendering observational methods inadequate for detecting true causal reliance. Moreover, in multi-agent systems, skill influence can persist independently of its original source, challenging the validity of current evaluation paradigms.
📝 Abstract
Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's own attribution. These signals show what the agent appears to use, not whether the skill changed its decision. We ask whether skill-augmented agents exhibit a \textbf{Reasoning Backroom}, a systematic gap between stated skill use and intervention-measured influence. We introduce BACKTRACE, an evaluation framework that pairs each skill-conditioned answer with a matched no-skill counterfactual, intervenes on skill meaning, wording, identity, content, and assignment, and elicits attribution only after the answer is committed. We instantiate the framework as BACKROOMBench, a verified testbed spanning controlled logic and competition mathematics, multiple skill conditions, single-agent and multi-agent settings, and diverse model families. Our evaluation reveals a pervasive provenance failure. Across models and domains, stated skill use often remains stable while causal reliance and signed utility vary, producing both silent uptake and performative use. Behavioral effects follow procedural content more reliably than displayed skill identity, whereas stated attributions respond strongly to artifact availability. Observational detectors based on direct skill-use claims, text mentions, trace similarity, and an LLM judge do not identify which decisions actually depend on the skill. In multi-agent systems, skill influence can survive communication even after its source is lost, while no-skill teams still name skills and sources that were never supplied. These findings establish the Reasoning Backroom as a general AI provenance problem whose audit requires intervention.