🤖 AI Summary
This study addresses the vulnerability of self-improving agents in open-ended tasks to reward hacking and capability blind spots due to the absence of reliable verifiers. To overcome this, it proposes a co-evolutionary framework wherein verifiers evolve alongside agents, optimizing for consistency rather than downstream scores. The approach synthesizes testable expressions through failure case clustering and implements gated selection via anchored reference sets and unlabeled consensus mechanisms, complemented by Double Ratchet lifecycle management. This work reveals the limitations of scoring-based authentication for self-evolving verifiers and highlights the critical role of anchors. Evaluated on MBPP+, the method improves consistency by 0.21 over handcrafted rules while retaining 88%–110% of performance gains, effectively mitigating rule gaming and repairing deficiencies.
📝 Abstract
We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.