False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "co-cheating" problem in self-evolving search agents, where proposers and solvers share errors, leading to inflated internal rewards while external accuracy stagnates. To mitigate this failure mode, we propose CrossFit, a mechanism that disrupts spurious consensus within feedback loops through source document cross-grouping and auxiliary solver verification, combined with a multi-sampling verification strategy. Experimental results demonstrate that CrossFit reduces the pseudo-label spurious agreement rate to below 3% and achieves average improvements of 8.4–8.8 points over standard self-evolution across seven downstream search benchmarks. By effectively breaking the homologous self-verification closed loop, this work provides a robust solution for enhancing the reliability and generalization of self-evolving search agents.
📝 Abstract
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
Problem

Research questions and friction points this paper is trying to address.

self-evolving search agents
co-cheating
pseudo-label correctness
false agreement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-evolving search agents
Co-cheating
CrossFit
Multi-sample verification
Pseudo-label correctness
M
Meijia Chen
Rutgers University
H
Hao Li
Independent Researcher
Z
Zheng Lu
Independent Researcher
H
Hongshan Lin
Independent Researcher
J
Junbai Tian
Independent Researcher
Y
Yichen Liu
University of California, San Diego
Z
Zijun Tian
Independent Researcher
Y
Yufan Zou
Independent Researcher
S
Shuhan Sun
Independent Researcher
H
Hankxin Chen
University of California, San Diego
Z
Zeyu Zhang
Independent Researcher
W
Weizhi Du
University of Michigan
Y
Yueting Li
Independent Researcher
Tianyu Shi
Tianyu Shi
University of Toronto
Reinforcement learningIntelligent Transportation SystemLarge Language ModelsAILLM agent
Alaa Khamis
Alaa Khamis
Associate Professor at KFUPM • Ex-GM Tech Leader
Smart mobilityconnected and automated vehiclesmachine learningcombinatorial optimization