🤖 AI Summary
This paper reveals that large language model (LLM)-based agent guardrails are vulnerable to "option channel attacks," resulting in critical safety failures. Through adversarial evaluation of seven open-source models using prompt injection, jailbreak testing, and synthetic tool-call logs, this study systematically examines their security as decision-making components. Results demonstrate that merely manipulating option labels or injecting irrelevant logs substantially degrades judgment accuracy, elevating pass-through rates to 100% and compromising all evaluated defenses, whereas deterministic rules achieve complete interception. To our knowledge, this work is the first to formalize this attack paradigm, establishing that type-based decision models should not serve as the ultimate safety backstop for LLM agents.
📝 Abstract
A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text. Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it. We evaluate seven open-weight models in that role and report the two error directions separately: a fail-open error allows a prohibited action and is a vulnerability; a fail-closed error blocks a permitted one and is only a cost. On prompt-injection, jailbreak and toxic-content screening, accuracy at the allow-or-block decision ranges from 36% to 72% against a chance level of 50%. A low error rate in one direction only reflects which answer a model defaults to: one allows nearly everything, another blocks nearly everything. On a synthetic suite of agent tool calls, six lines of server log text that say nothing about the policy raise a gate's fail-open rate from 0% to 63% on a policy it otherwise decides correctly. Giving the permissive option a misleading name, with its definition and the judged text untouched, raises that rate to between 93% and 100% on the four models that place the label in their input. Every defense we tested is defeated, either by an attacker who targets its mechanism or by attacker-controlled text. Escalating the least confident decisions does not help either: a decision an attack has reversed is no less confident than the one it replaced. Parsing each policy field into a typed value does eliminate one attack, but it also makes the model unnecessary: a deterministic rule over those values reaches 100% accuracy on all six policies. These models can reduce how many cases reach a reviewer, but on this evidence they should not be the component that decides. Code is available at https://github.com/ArminAzizi98/option-channel-attack.