ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents

📅 2026-05-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of AI agents predominantly focus on static outputs, failing to uncover behavioral flaws that emerge during multi-turn interactions or under adversarial or high-pressure conditions. This work proposes a scalable and auditable dynamic evaluation infrastructure that shifts assessment from a single-score paradigm to an evidence-based, process-oriented analysis. By integrating adversarial multi-turn testing, turn-level behavioral trajectory tracing, multi-reviewer consensus scoring, and evidence-linked reporting mechanisms, the framework enables comprehensive scrutiny of agent behavior. It supports flexible expansion across evaluation dimensions and effectively exposes vulnerabilities in otherwise high-performing agents across diverse domains—including customer service, medical triage, privacy-sensitive scenarios, and code generation. Notably, experiments demonstrate that even small, quantized local LLMs can serve as efficient challengers capable of rigorously evaluating production-grade agents powered by state-of-the-art large language models.
📝 Abstract
AI agents are entering high-risk production settings, where they use tools, retain context, follow policies, handle private data, and interact with users over multiple turns. Yet many evaluation methods still judge isolated outputs or static tasks, missing failures that emerge through trajectory, pressure, and adversarial interaction. We introduce ProofAgent Harness, open infrastructure for scalable, auditable, and adversarial AI agent evaluation. The harness provides evaluation infrastructure around an agent: it curates evaluation intelligence, runs adversarial multi-turn trials, captures behavioral traces, applies post-hoc multi-juror scoring, resolves disagreement, and produces evidence-linked reports. Its open design allows developers and researchers to extend domains, traps, metrics, juror personas, scoring rules, and reporting formats. At its core is Adversarial Multi-Juror Scoring with Turn-Level Audit, which evaluates completed agent behavior under pressure using calibrated juror personas, consensus checks, and turn-level evidence. Experiments across customer support, medical triage, privacy and security, and code generation agents show that strong agents fail selectively through weak metrics, fragile turns, unsafe reframing, and manipulation paths. We also find that a small quantized local Harness LLM can challenge production agents powered by best-in-class large LLMs, suggesting that evaluation capability emerges from the full harness pipeline rather than model scale alone. ProofAgent Harness turns AI agent evaluation from a static score into scalable adversarial evaluation infrastructure: repeatable, evidence-backed, extensible, and actionable before deployment.
Problem

Research questions and friction points this paper is trying to address.

adversarial evaluation
AI agents
multi-turn interaction
behavioral failure
high-risk deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Evaluation
Multi-Juror Scoring
Turn-Level Audit
AI Agent Testing
Open Evaluation Infrastructure
🔎 Similar Papers
No similar papers found.
F
Fouad Bousetouane
ProofAgent.ai, The University of Chicago, USA