Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of tool invocations by LLM agents to tampering during transmission, which leads to intent-deviating executions that are difficult to detect. To investigate this, we introduce the concept of Intent-Execution Correspondence (IEC) and construct IEC-Bench, the first benchmark to quantitatively measure the effects of in-transit modifications. We further propose IntAct, a protocol that ensures invocation integrity or enables secure refusal through execution-free observation and tamper-proof delivery mechanisms. Our evaluation reveals that 12% of invocations are tampered with, predominantly resulting in silent failures, and highlights inherent biases in conventional attribution methods. IntAct successfully recovers 79.2% of these failures while significantly reducing retry overhead, thereby providing an effective end-to-end remediation framework for securing LLM agent tool calling.
📝 Abstract
Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call without notice. When the changed call fails, the agent retries a correct call, which costs users time and money. Benchmarks and failure analyses do not see the change, because they read the call and its result but not what a hop received. We define intent-execution correspondence (IEC) as the property that the executed action matches the action the emitted call denotes under the tool contract. Our protocol observes what each hop received without executing the call, and names the first hop that changed it by the receiver's own parser. IntAct then delivers the call in a form that this hop cannot alter, or refuses the call. We build IEC-Bench from the changes observed in real-world use, with chains of dependent calls under the execution paths of 4 widely-used harnesses. In 47,828 shell calls within production sessions, Claude Code's Bash tool changes 12.0% of the calls that carry code, escape sequences, or long text. For 80.7% of the calls whose backslashes are changed, the wrong action runs without any reported error. All 10 measured harnesses change a call. Trajectory-based judgment attributes 95.1% of the production failures to the LLM, although the path caused more than half of them. On IEC-Bench, the path raises the token cost per passed task 2.4 times (up to 12.3 times). A hop that changes a call also hides the changes after it, so 55.1% of the failures on one path appear only after its first hop is repaired. IntAct, deployed in a commercial product, recovers 79.2% of the failures with a changed call. Harnesses should therefore be designed and tested hop-by-hop to ensure a correct call executes as intended or is refused.
Problem

Research questions and friction points this paper is trying to address.

LLM Agents
Tool Calls
Intent-Execution Correspondence
Fault Attribution
Execution Harnesses
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intent-Execution Correspondence
Tool Call Mutation
LLM Agents
IntAct
IEC-Bench
🔎 Similar Papers