Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of self-improving agents in open-ended tasks to reward hacking and capability blind spots due to the absence of reliable verifiers. To overcome this, it proposes a co-evolutionary framework wherein verifiers evolve alongside agents, optimizing for consistency rather than downstream scores. The approach synthesizes testable expressions through failure case clustering and implements gated selection via anchored reference sets and unlabeled consensus mechanisms, complemented by Double Ratchet lifecycle management. This work reveals the limitations of scoring-based authentication for self-evolving verifiers and highlights the critical role of anchors. Evaluated on MBPP+, the method improves consistency by 0.21 over handcrafted rules while retaining 88%–110% of performance gains, effectively mitigating rule gaming and repairing deficiencies.
📝 Abstract
We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.
Problem

Research questions and friction points this paper is trying to address.

self-improving agents
verifier reliability
reward hacking
open-ended tasks
co-evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

co-evolving verifier
inspectable graders
reward hacking prevention
anchored reference set
Double Ratchet
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xing Zhang
AWS Forward Deployed Engineering
G
Guanghui Wang
AWS Forward Deployed Engineering
Y
Yanwei Cui
AWS Forward Deployed Engineering
Ziyuan Li
Ziyuan Li
Associate Professor, School of Optics and Photonics, Beijing Institute of Technology
Optoelectronicssemiconductornanowireplasmonicsoptical antennas
W
Wei Qiu
HSBC Holdings Plc., HSBC Technology Center, China
B
Bing Zhu
HSBC Holdings Plc., HSBC Technology Center, China
P
Peiyang He
AWS Forward Deployed Engineering