Score
Design, build, and evaluate algorithms and tools that identify and filter contaminated records in datasets and corpora—producing binary clean/contaminated classifications, overlap/contamination scores, or filters to remove or flag items that appear in training or external sources. Also develop auditing procedures and metrics to detect pretraining exposure and memorization (e.g., unusually fast loss reduction or backbone movement), set admission thresholds for external data, and improve the reliability of held-out evaluation.
This paper addresses data contamination in large language model (LLM) evaluation—unintended overlap between training data and evaluation benchmarks that inflates performance metrics and misleads generalization assessment. Methodologically, it introduces the first taxonomy of contamination detection based on model information dependency: white-box (leveraging gradients or memory traces), gray-box (assessing output consistency), and black-box approaches. It further proposes a novel paradigm of dynamic benchmark construction and LLM-driven contamination-free evaluation, complemented by an integrated mitigation strategy comprising data updating, rewriting, and prevention. The contributions are threefold: (1) exposing systemic vulnerabilities in current LLM evaluation practices; (2) establishing the first comprehensive, classification-based governance framework spanning contamination detection, unbiased evaluation, and proactive prevention; and (3) providing both theoretical foundations and practical guidelines for developing more rigorous and trustworthy LLM evaluation protocols.
LLM evaluation is frequently compromised by train-test data contamination, yet existing contamination detection methods rely on unvalidated underlying assumptions. This paper systematically reviews 50 studies to propose the first taxonomy of assumptions in contamination detection, identifying eight core assumption categories and empirically evaluating three representative ones. We introduce a novel hypothesis-driven evaluation paradigm integrating membership inference attacks (MIAs), distributional shift analysis, and large-scale contamination simulation. Our findings reveal that state-of-the-art MIAs perform near-randomly on real LLM pretraining data; distributional shifts substantially degrade detection reliability; and LLMs preferentially learn statistical patterns over memorizing individual instances. These results challenge the default “instance memorization” assumption, offering both theoretical foundations and methodological guidance for trustworthy LLM evaluation.
Training data opacity—particularly in closed-source LLMs—leads to evaluation data contamination, severely undermining the reliability of performance assessment and hindering practical progress in NLP. Method: We systematically survey over 50 contamination detection studies and propose the first unified taxonomy spanning output statistics, gradient/activation tracing, training set reconstruction, prompt perturbation, and counterfactual reasoning—clarifying applicability boundaries and inherent limitations of each approach. We advocate integrating contamination detection into standard evaluation pipelines and formally characterize how training data leakage induces evaluation distortion. Contribution/Results: Our work enables systematic modeling and mitigation of contamination bias. It provides actionable guidelines for industry practitioners to conduct trustworthy model evaluation and supports academic efforts in establishing fair, contamination-aware benchmarks—thereby advancing both methodological rigor and real-world deployment integrity in LLM assessment.
Pretraining data filtering strategies intended to reduce harmful content inadvertently exacerbate representational underrepresentation of marginalized groups, thereby amplifying demographic bias at the data level. Method: We systematically reviewed 55 English-language large language model technical reports to construct the first integrated data governance evaluation framework balancing safety and fairness. Through controlled experiments and quantitative bias analysis across mainstream filtering strategies, we measured their impact on group-level representation. Contribution/Results: Our analysis reveals that such strategies reduce text associated with disadvantaged groups by 12.7%–38.4% on average—significantly worsening representational disparity. This study provides the first empirical evidence refuting the “safety implies fairness” assumption in AI governance. We propose a co-optimization paradigm that jointly addresses content safety and equitable group representation, advocating for fairness-aware data curation in foundation model development.
This work exposes a critical vulnerability in benchmark contamination detection for reasoning models (LRMs): developers can substantially inflate leaderboard scores by injecting evaluation data into supervised fine-tuning or RL stages (e.g., PPO/GRPO), while existing contamination detection methods fail almost entirely. The authors systematically analyze two realistic contamination scenarios, combining theoretical analysis with empirical validation to demonstrate that objective clipping in PPO-style algorithms and chain-of-thought fine-tuning effectively obscure contamination signals—degrading mainstream detectors’ performance to chance level. The core contribution is the first identification of LRM-specific contamination concealment mechanisms, empirically confirming a fundamental flaw in current detection paradigms. This work provides both a crucial caution and technical foundation for establishing trustworthy LRM evaluation frameworks.
This paper addresses the challenge of detecting training data contamination in large language models (LLMs). We propose the Data Contamination Quiz (DCQ), a parameter-free, zero-shot contamination detection method that frames detection as a multiple-choice discrimination task. DCQ generates semantically preserved, word-level perturbations to construct distractors and quantifies contamination likelihood by measuring an LLM’s preference strength for the original input over perturbed alternatives. Our approach introduces a novel “semantically indistinguishable yet lexically matchable” paradigm—requiring no training data or model fine-tuning—and inherently bypasses copyright-aware safety filters. To enhance robustness, we incorporate positional bias correction. Evaluated across multiple LLMs and datasets, DCQ achieves state-of-the-art performance. Moreover, it systematically reveals, for the first time, significantly higher levels of memorized data contamination in LLMs than previously recognized.
This work addresses the challenge of detecting training data contamination in large language models (LLMs) without access to the original training corpus. We propose CoDeC, a contamination detection method that analyzes shifts in model confidence during in-context learning (ICL): specifically, it identifies memory traces by measuring anomalous confidence degradation when models process samples from their own training set—contrasted with unseen data. Crucially, we observe and formalize—for the first time—that training data interferes with ICL dynamics, inducing statistically detectable confidence suppression. This effect serves as an interpretable, model-agnostic contamination signal. Extensive experiments across multiple open-weight LLMs demonstrate that CoDeC’s contamination scores robustly separate seen (contaminated) from unseen (clean) data. The method is fully automated, requires no training data or model fine-tuning, and integrates seamlessly into standard evaluation pipelines. CoDeC establishes a new paradigm for privacy-aware data auditing and trustworthy model assessment.
This paper addresses the problem of overconfident evaluation of large language models (LLMs) due to training data contamination—particularly in short-text domains such as psychological scales. We propose LogProber, the first quantifiable contamination detection method tailored for short texts. It leverages sentence-level token log-probability modeling, sequence likelihood estimation, and statistical significance testing to efficiently identify data leakage in low-resource settings. Key contributions include: (i) the first empirical demonstration that instruction tuning and similar training paradigms can induce “stealth contamination”—where contamination exists despite negligible changes in token probabilities; and (ii) a systematic characterization of detection capability boundaries and failure conditions across mainstream training paradigms. Evaluated on diverse psychological questionnaire datasets, LogProber achieves high-precision contamination identification, offering a reproducible, interpretable diagnostic tool for fair LLM evaluation.
This study addresses a critical gap in data cleaning research—the lack of large-scale, real-world dirty datasets that hinder the effective evaluation of methods in practical settings. To bridge this gap, the authors construct the first large-scale dirty dataset comprising real postal addresses paired with their ground-truth counterparts. Leveraging this dataset, they conduct a systematic benchmarking study of state-of-the-art data cleaning approaches. Their experiments reveal substantial limitations of current methods when applied to real-world data, underscoring the need for more robust and context-aware techniques. The dataset and empirical findings not only establish a reliable benchmark for future research but also provide actionable insights to guide the development of data cleaning solutions tailored to real-world applications.
Existing contamination detection methods rely on assumptions such as access to training data, handcrafted statistics, or predefined labels, limiting their applicability in real-world settings. This work proposes a novel approach that dispenses with such assumptions by constructing depth profiles via linear probes in residual streams, introducing a metric termed “excess separability,” and combining label permutation tests, item-wise bootstrapping, and size-matched placebo control sets to detect whether a model has been exposed to test data while controlling for confounding factors. The method achieves a substantially reduced false positive rate—down to 0.02—rejects fragile ablations, and demonstrates effectiveness on real Transformer models: it finds no evidence of contamination across four Pile subsets. All implementation and auditing code is publicly released.
This work addresses a critical yet overlooked issue in adaptive data cleaning: fluctuating sample removal counts caused by varying partition granularities introduce budget confounding bias, leading to spurious performance gains falsely attributed to contamination identification capability. To enable fair evaluation, the authors propose an operating-point-matched assessment framework that aligns removal budgets with recall rates and incorporates threshold-agnostic metrics (AUROC and AUPRC). They systematically uncover and resolve this budget confounding problem for the first time, introducing a multi-cue adaptive cleaner—integrating learning difficulty reweighting, Euclidean distance guidance, and fine-grained partitioning—and a false positive decomposition analysis. Experiments on CIFAR-10 and ImageNet-100 reveal that most existing methods lose their apparent advantage under matched operating points, demonstrating genuine efficacy only under low contamination rates or high-recall, heavily corrupted scenarios.