Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether self-reflective signals in reasoning models genuinely reflect capability and effectively guide tool-calling decisions. By injecting first-person expressions of confidence or doubt into reasoning traces as counterfactual interventions, this work introduces the concept of "steerability" to decouple signal sensitivity from goal-directedness, enabling the quantification of endogenous model mechanisms without post-training. Experiments across multiple families of open-source reasoning models reveal that while models are highly sensitive to confidence signals, their directional capability remains remarkably poor, performing only marginally better than random baselines. These findings demonstrate that current reasoning models lack genuine capability-tracking mechanisms, offering a novel paradigm for understanding and evaluating their introspective reliability.
📝 Abstract
Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to the model (numerical prediction, confidence token, prompt injection), and whether it changes the model's subsequent action. We isolate the third question. At a fixed point in otherwise identical reasoning trajectories, we insert a single first-person sentence expressing either confidence or doubt; the model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation. We call this behavioral response Nudgeability and measure it along two dimensions: sensitivity, how strongly confidence and doubt change delegation rates, and targeting, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points, and 53 to 70 points for the larger provider-served models. This responsiveness is poorly targeted: a median 42% of induced flips are well-targeted, only a +2 percentage-point lift over a random-selection baseline. Confidence language is thus a strong control surface for delegation, but current models use it only weakly in accordance with their actual competence. Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature.
Problem

Research questions and friction points this paper is trying to address.

reasoning language models
tool delegation
self-reflection
nudgeability
confidence signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nudgeability
Causal Inference
Self-Reflection
Tool Delegation
Reasoning Models