Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether errors in deterministic rules and large language model (LLM) judgment layers propagate independently during runtime action gating for AI agents. Evaluating multi-layer gating combinations across 1,119 annotated actions, the authors propose a Multiplicative Equivalent Layer metric to quantify individual layer contributions, complemented by statistical testing and dual-criteria scoring. The findings reveal that single-layer accuracy cannot reliably predict combined performance gains. Notably, hybrid rule-LLM architectures exhibit significantly lower coupling than purely LLM-based ensembles, achieving approximately twofold gains compared to merely 1.2–1.4× for the latter. Furthermore, this work highlights critical risks associated with unidentifiable difficulty attribution and model version drift. Overall, these insights provide empirical guidance for designing robust, decoupled gating mechanisms in agent systems.
📝 Abstract
Runtime gates for agent tool calls are stacked on the assumption that their errors multiply. We test it on 1,119 labelled agent actions from three corpora, without an adaptive adversary. The stack has one deterministic rule layer and four LLM judges, three of them re-collected with the served model recorded on every call. We read each stack as a number of multiplication-equivalent layers, n_mult, with its floor under perfect coupling. Under the STRICT miss definition (escalation to a human scored as not stopped), any two judges compose to about 1.2 to 1.4 layers (φ median +0.430, 6 of 6 pairs significant, floors 1.02 to 1.17). The rule layer plus one judge composes to 1.86 to 2.09 layers (φ median +0.014, 0 of 4 significant, floors 1.01 to 1.09). Under PRIMARY (escalation scored as caught) the bands are 1.21 to 1.57 and 1.80 to 2.13. Intervals separate on the pooled data, point estimates split on each corpus, and a third-vendor judge lands in the judge band. Solo accuracy does not predict what a layer adds: a cloud rule pack lowers the rule layer's solo miss rate by 20% and adds no new joint coverage. The difficulty share of judge coupling is not identifiable: 31.8% to 61.8% depending on the probe and the miss definition. One judge tier was served by an unrequested model version in 50 of 112 batches, concentrated on the external corpus. That event overturned a pre-declared analysis rule, and the scoring of review verdicts reversed five conclusions. We report both.
Problem

Research questions and friction points this paper is trying to address.

agent safety gates
error independence
layer stacking
runtime evaluation
LLM judges
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Tool Calls
Runtime Gates
LLM Judges
Error Coupling
Safety Evaluation
🔎 Similar Papers
No similar papers found.