🤖 AI Summary
This study addresses the challenge that the reliability of AI judges evaluating mathematical reasoning remains difficult to verify, with uninterpretable failure modes. To overcome the limitations of conventional benchmarking, this work proposes a two-stage surrogate-guided framework that leverages a multi-agent system—orchestrating Codex and Claude Code—to perform adversarial error injection, subsequently extracting generalizable failure mechanisms from these attacks via policy distillation. The contributions of this research lie in systematically exposing evaluation vulnerabilities in frontier judge models such as GPT-5.6, successfully circumventing their detection mechanisms. Furthermore, it reveals that error detection rates for research-level texts are significantly lower than those for foundational mathematical proofs. These findings establish a novel paradigm for enhancing the robustness of AI judges.
📝 Abstract
As AI agents surpass human performance, it becomes exceedingly hard for system designers to evaluate them directly and understand their failure modes. Consequently, agents themselves are being deployed extensively to evaluate, judge, and provide feedback on model traces. But this raises an important question: how can we trust the judge? In this work, we propose an agent-guided method to find weaknesses of agentic judges that expose interpretable failure mechanisms. Our method focuses on mathematical reasoning and proceeds in two stages: first, we deploy adversarial agents to mutate a set of sound proofs by introducing errors, attempting to misguide judges---in other words, injecting errors that judges are unable to catch. Then, we distill these attempts into a small set of mutation strategies which allow us to analyze the failure modes of the judges. To ensure that these strategies are not overfit to the initial set of proofs, we evaluate them by applying the mutation strategies to a held-out set of proofs and querying the same judge. We deploy our method on GPT-5.6-sol and Claude Opus 5, paired with their agent orchestrators, Codex and Claude Code, respectively. These are used both as mutators to introduce errors and as judges to evaluate correctness of mathematical reasoning. We find that across all the agentic judges, we are able to distill mutation strategies that consistently bypass their evaluations, thereby enabling us to ascertain actionable failure modes. Our analysis also reveals that judge reliability degrades at the frontier: errors in Olympiad-level proofs or graduate-level mathematical texts are detected more consistently, whereas flaws in research-level manuscripts are more likely to escape detection.