Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deployment costs and performance misreporting associated with System-1 decision models in LLM agents. To this end, it constructs a paired evaluation framework encompassing eleven decision points and 7,283 cases, introducing a self-auditing mechanism to rectify pipeline errors. By integrating byte-level consistency verification, zero-shot routing evaluation, and RAG gating techniques, the work achieves full-process transparency. The findings reveal that although the Jev model outperforms Laya on nine of eleven tasks, neither surpasses a random baseline. Following error correction, the actual savings amount to only 4.3%, exposing severe design confounding and data leakage risks. Ultimately, this research establishes a rigorous benchmark for assessing the genuine efficacy of System-1 models.
📝 Abstract
Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya changes 30% of its answers when the option order is reversed and degrades sharply with many or similar candidates (31% at 50 nearest-neighbour tools, vs. 98% for Jev on items with a unique correct tool). We also audit our own pipeline. Three analysis errors and one design confound distorted headline deployment claims: an omitted pre-screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end-to-end quality (58% vs. 98%), in-sample thresholds (5% target, up to 17% held-out misses), and a "channel effect" on injection false positives that vanishes with channel-native content. Two other suspected confounds did not change the conclusions. All cases, raw outputs and analysis code are available at https://github.com/David-DL-Space/sys1-eval.
Problem

Research questions and friction points this paper is trying to address.

System-1 decision models
LLM agent harnesses
model evaluation
robustness
pipeline auditing
Innovation

Methods, ideas, or system contributions that make the work stand out.

System-1 decision models
Paired evaluation
LLM agent harnesses
Self-auditing
Robustness testing
🔎 Similar Papers
No similar papers found.