On The Fragility of Benchmark Contamination Detection in Reasoning Models

📅 2025-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work exposes a critical vulnerability in benchmark contamination detection for reasoning models (LRMs): developers can substantially inflate leaderboard scores by injecting evaluation data into supervised fine-tuning or RL stages (e.g., PPO/GRPO), while existing contamination detection methods fail almost entirely. The authors systematically analyze two realistic contamination scenarios, combining theoretical analysis with empirical validation to demonstrate that objective clipping in PPO-style algorithms and chain-of-thought fine-tuning effectively obscure contamination signals—degrading mainstream detectors’ performance to chance level. The core contribution is the first identification of LRM-specific contamination concealment mechanisms, empirically confirming a fundamental flaw in current detection paradigms. This work provides both a crucial caution and technical foundation for establishing trustworthy LRM evaluation frameworks.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessComputer Vision: Adversarial Attacks & Robustness

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
Leaderboards for LRMs have turned evaluation into a competition, incentivizing developers to optimize directly on benchmark suites. A shortcut to achieving higher rankings is to incorporate evaluation benchmarks into the training data, thereby yielding inflated performance, known as benchmark contamination. Surprisingly, our studies find that evading contamination detections for LRMs is alarmingly easy. We focus on the two scenarios where contamination may occur in practice: (I) when the base model evolves into LRM via SFT and RL, we find that contamination during SFT can be originally identified by contamination detection methods. Yet, even a brief GRPO training can markedly conceal contamination signals that most detection methods rely on. Further empirical experiments and theoretical analysis indicate that PPO style importance sampling and clipping objectives are the root cause of this detection concealment, indicating that a broad class of RL methods may inherently exhibit similar concealment capability; (II) when SFT contamination with CoT is applied to advanced LRMs as the final stage, most contamination detection methods perform near random guesses. Without exposure to non-members, contaminated LRMs would still have more confidence when responding to those unseen samples that share similar distributions to the training set, and thus, evade existing memorization-based detection methods. Together, our findings reveal the unique vulnerability of LRMs evaluations: Model developers could easily contaminate LRMs to achieve inflated leaderboards performance while leaving minimal traces of contamination, thereby strongly undermining the fairness of evaluation and threatening the integrity of public leaderboards. This underscores the urgent need for advanced contamination detection methods and trustworthy evaluation protocols tailored to LRMs.
Problem

Research questions and friction points this paper is trying to address.

Benchmark contamination detection is alarmingly easy to evade in reasoning models
RL training methods inherently conceal contamination signals from most detection approaches
Contaminated models achieve inflated leaderboard performance while leaving minimal detection traces
Innovation

Methods, ideas, or system contributions that make the work stand out.

RL training conceals contamination detection signals
PPO objectives cause detection concealment in models
Contaminated models evade memorization-based detection methods
🔎 Similar Papers
2024-06-26Conference on Empirical Methods in Natural Language ProcessingCitations: 0
H
Han Wang
University of Illinois Urbana-Champaign
H
Haoyu Li
University of Illinois Urbana-Champaign
B
Brian Ko
University of Washington
H
Huan Zhang
University of Illinois Urbana-Champaign