Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过控制令牌注入方法,揭示了工具使用型语言模型的安全性是模型和其解码框架共同属性,并提出了解决方案如输入净化、解析器加固等。
📝 Abstract
The safety of a tool-using language model agent is usually treated as a property of the model alone. We give controlled, full-precision evidence that it is instead a joint property of the model and the software that renders its chat template and parses its tool calls, the decoding harness, and that both halves are attackable from untrusted input. On the released gpt-oss-20b reasoning model under its published tool sandbox, appending a single string of the model's own channel-control tokens to a user message makes the tokenizer render a reasoning turn that is already complete, so the model writes no chain-of-thought and proceeds directly to the tool call. Across forty tasks the model already completes, the reasoning channel falls from a mean of 52.5 tokens to zero on every trial while the http.post still fires on every trial. A rule monitor and a cross-family language-model monitor detect the unsafe request on all plain trials and no forged trials, and on overtly malicious requests the attack converts 39.6% of the model's refusals into completed exfiltrations. Separately, whether an identical tool-call generation fires is decided by the harness parser, not the model: a truncation-tolerant regular expression fires a call whose closing token is missing while a strict one drops it, and two parsers shipped for the Gemma agent give opposite outcomes on identical greedy generations, firing on all twenty-four trials and on none. We show the suppression can be delivered indirectly and characterize its dependence on the chat template across two more reasoning models, and we evaluate input sanitization, parser hardening, and empty-reasoning detection as defenses; flagging an absent trace catches the basic attack but not an adaptive benign decoy. All measurements use greedy decoding on publicly released models. Code and per-trial logs: https://github.com/Usama1002/deleting-the-trace
Problem

Research questions and friction points this paper is trying to address.

tool-using language model
chain-of-thought
reasoning-based oversight
control-token injection
decoding harness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Control-Token Injection
Chain-of-Thought Suppression
Decoding Harness Vulnerability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Muhammad Usama
Braindeck Inc, Seoul, Republic of Korea
K
Khair Un Nisa
University of Wah, Pakistan
S
Summer Yeoreum Jung
Braindeck Inc, Seoul, Republic of Korea