One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper reveals that large language model (LLM)-based agent guardrails are vulnerable to "option channel attacks," resulting in critical safety failures. Through adversarial evaluation of seven open-source models using prompt injection, jailbreak testing, and synthetic tool-call logs, this study systematically examines their security as decision-making components. Results demonstrate that merely manipulating option labels or injecting irrelevant logs substantially degrades judgment accuracy, elevating pass-through rates to 100% and compromising all evaluated defenses, whereas deterministic rules achieve complete interception. To our knowledge, this work is the first to formalize this attack paradigm, establishing that type-based decision models should not serve as the ultimate safety backstop for LLM agents.
📝 Abstract
A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text. Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it. We evaluate seven open-weight models in that role and report the two error directions separately: a fail-open error allows a prohibited action and is a vulnerability; a fail-closed error blocks a permitted one and is only a cost. On prompt-injection, jailbreak and toxic-content screening, accuracy at the allow-or-block decision ranges from 36% to 72% against a chance level of 50%. A low error rate in one direction only reflects which answer a model defaults to: one allows nearly everything, another blocks nearly everything. On a synthetic suite of agent tool calls, six lines of server log text that say nothing about the policy raise a gate's fail-open rate from 0% to 63% on a policy it otherwise decides correctly. Giving the permissive option a misleading name, with its definition and the judged text untouched, raises that rate to between 93% and 100% on the four models that place the label in their input. Every defense we tested is defeated, either by an attacker who targets its mechanism or by attacker-controlled text. Escalating the least confident decisions does not help either: a decision an attack has reversed is no less confident than the one it replaced. Parsing each policy field into a typed value does eliminate one attack, but it also makes the model unnecessary: a deterministic rule over those values reaches 100% accuracy on all six policies. These models can reduce how many cases reach a reviewer, but on this evidence they should not be the component that decides. Code is available at https://github.com/ArminAzizi98/option-channel-attack.
Problem

Research questions and friction points this paper is trying to address.

Agent Guardrails
Typed Decision Models
Option-Channel Attack
Prompt Injection
AI Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Option-Channel Attack
Typed Decision Models
Agent Guardrails
Prompt Injection
Fail-open Vulnerability