Tool Mediation Alters Refusal Mechanisms in Large Language Models

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unclear mechanisms and degraded robustness of large language models (LLMs) in refusing harmful requests during tool interactions. By integrating representation geometry, neuron-level analysis, and intervention experiments, we investigate how tool mediation alters the models' internal representations and computational distributions. Our findings reveal that tool interaction does not diminish the models' capacity for harm perception; rather, it elevates the effective refusal threshold and introduces novel vulnerabilities. This work elucidates the fundamental reasons why conventional safety evaluations cannot be directly transferred to agent-based scenarios, demonstrating that tool-augmented environments intrinsically compromise the models' resistance to harmful inputs. Ultimately, these insights offer a new perspective on the safety alignment of autonomous agents.
πŸ“ Abstract
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms across a diverse set of open-weight language models. We find that information about the harmfulness of a request remains strongly encoded in the model's representations and transfers across conversational and tool-mediated inputs. Evidence from representation geometry and neuron-level analysis further indicates that the two interaction modes systematically distribute harm-related computation differently. Crucially, while conversational inputs can be refused at relatively low levels of perceived harmfulness, tool-mediated inputs remain permissive until harmfulness crosses a substantially higher effective refusal threshold. Moreover, tool-mediated refusal is also more brittle: progressively weakening the refusal computation disrupts tool-mediated refusal at lower intervention strengths than conversational refusal, even when benign capabilities remain intact. Together, our findings indicate that tool mediation does not simply reduce the internal perception of harm, but instead impacts its conversion into refusal. Overall, this suggests tool-mediated environments may intrinsically reduce robustness of models to harmful requests, and that conventional safety evaluations may not fully transfer to LLM agents.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Tool Mediation
Refusal Mechanisms
AI Safety
Harmful Requests
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tool Mediation
Refusal Mechanisms
Representation Geometry
Neuron-level Analysis
Refusal Threshold