When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the latent degradation and evaluation blind spots induced by internal activation steering in tool-calling scenarios of large language models (LLMs). We propose SAKIKO, an auditing framework that systematically evaluates the genuine remediation effects of internal interventions through directional error discovery, channel-keyed intervention, and target-parsing verification. Furthermore, this work pioneers an outcome-parsing-based adjudication mechanism and forward-freezing statistical licensing, revealing that behavioral modification does not equate to fundamental repair. Experiments across seven LLMs demonstrate that most existing interventions incur severe side effects, thereby establishing the necessity of outcome-level adjudication for ensuring intervention rigor.
📝 Abstract
Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: https://github.com/ruizheliUOA/mechanistic-tool-use-llm.
Problem

Research questions and friction points this paper is trying to address.

tool-using LLMs
mechanistic auditing
internal interventions
representation repair
activation steering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mechanistic Auditing
Representation Repair
Activation Steering
Tool-Using LLMs
Destination Auditing
🔎 Similar Papers
J
Jiayi Li
University of the Chinese Academy of Sciences, China
R
Ruizhe Li
School of Computer Science, University of Birmingham, UK