VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of output verification for LLM agents in long-horizon tasks, where reference standards are typically unavailable. We propose a reference-free self-improving verification framework that transforms LLMs into intelligent verifiers equipped with tool-calling capabilities and reusable skills. By constructing workspaces through repeated sampling, the framework refines final outputs via divergence resolution, consensus challenging, and environmental evidence retrieval mechanisms. Our analysis reveals a novel insight: divergence exposes correctness, whereas consensus may conceal errors. Experimental results demonstrate that the proposed method achieves state-of-the-art best-of-N selection scores across five benchmarks, yielding performance improvements exceeding six points for two leading frontier models. Furthermore, we release an open-source dataset comprising 26,000 trajectories to facilitate future research.
📝 Abstract
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
Problem

Research questions and friction points this paper is trying to address.

agentic verification
long-horizon tasks
LLM agents
output verification
repeated sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Verification
Long-Horizon Tasks
Disagreement Resolver
Consensus Challenger
Self-Improving Skills