Score
Designs and implements metrics and procedures to quantify contamination in datasets or model outputs, producing per-item contamination scores and aggregated scalar measures (including weighting by contamination or cognitive type). Builds analysis pipelines and reports that identify, rank, and summarize high-impact contaminated items and compute dataset-level contamination statistics for downstream evaluation or mitigation.
This paper addresses data contamination in large language model (LLM) evaluation—unintended overlap between training data and evaluation benchmarks that inflates performance metrics and misleads generalization assessment. Methodologically, it introduces the first taxonomy of contamination detection based on model information dependency: white-box (leveraging gradients or memory traces), gray-box (assessing output consistency), and black-box approaches. It further proposes a novel paradigm of dynamic benchmark construction and LLM-driven contamination-free evaluation, complemented by an integrated mitigation strategy comprising data updating, rewriting, and prevention. The contributions are threefold: (1) exposing systemic vulnerabilities in current LLM evaluation practices; (2) establishing the first comprehensive, classification-based governance framework spanning contamination detection, unbiased evaluation, and proactive prevention; and (3) providing both theoretical foundations and practical guidelines for developing more rigorous and trustworthy LLM evaluation protocols.
LLM evaluation is frequently compromised by train-test data contamination, yet existing contamination detection methods rely on unvalidated underlying assumptions. This paper systematically reviews 50 studies to propose the first taxonomy of assumptions in contamination detection, identifying eight core assumption categories and empirically evaluating three representative ones. We introduce a novel hypothesis-driven evaluation paradigm integrating membership inference attacks (MIAs), distributional shift analysis, and large-scale contamination simulation. Our findings reveal that state-of-the-art MIAs perform near-randomly on real LLM pretraining data; distributional shifts substantially degrade detection reliability; and LLMs preferentially learn statistical patterns over memorizing individual instances. These results challenge the default “instance memorization” assumption, offering both theoretical foundations and methodological guidance for trustworthy LLM evaluation.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.
This work exposes a critical vulnerability in benchmark contamination detection for reasoning models (LRMs): developers can substantially inflate leaderboard scores by injecting evaluation data into supervised fine-tuning or RL stages (e.g., PPO/GRPO), while existing contamination detection methods fail almost entirely. The authors systematically analyze two realistic contamination scenarios, combining theoretical analysis with empirical validation to demonstrate that objective clipping in PPO-style algorithms and chain-of-thought fine-tuning effectively obscure contamination signals—degrading mainstream detectors’ performance to chance level. The core contribution is the first identification of LRM-specific contamination concealment mechanisms, empirically confirming a fundamental flaw in current detection paradigms. This work provides both a crucial caution and technical foundation for establishing trustworthy LRM evaluation frameworks.
This study systematically investigates how data contamination affects the performance evaluation of pre-trained language models (PLMs; e.g., RoBERTa, GPT-2) and large language models (LLMs; e.g., LLaMA, StarCoder) on code intelligence tasks—namely translation, generation, and summarization. We design controlled experiments across four contamination scenarios: input-only, output-only, unpaired, and paired contamination, covering both pretraining–finetuning–inference and direct inference paradigms. Key findings reveal that paired contamination induces significant performance inflation only in LLMs under direct inference or PLMs under minimal finetuning; unpaired and single-sided contamination exert negligible effects; and PLMs exhibit robustness under standard finetuning, whereas LLMs—relying heavily on contextual pairing—are more vulnerable. Crucially, this work provides the first empirical evidence that contamination does not necessarily lead to overestimation—a counterintuitive insight that challenges conventional evaluation assumptions. It establishes new benchmarks and practical guidelines for trustworthy evaluation and secure deployment of code models.
Existing synthetic data contamination detection methods rely solely on token-level overlap, failing to identify latent benchmark contamination—where no lexical repetition occurs but semantic similarity persists—thereby severely compromising model evaluation validity. To address this, we propose the first four-tiered contamination detection framework, spanning token-level, semantic, reasoning-pattern, and performance-collapse dimensions. Our approach integrates semantic embedding comparison, reasoning-path analysis, anomalous performance monitoring, and controlled experimental design to systematically uncover contamination across semantic and reasoning levels. Extensive experiments on MMLU, GSM8K, and HumanEval demonstrate that our method achieves an average F1-score of 0.76—outperforming the state-of-the-art by 26.5%—significantly enhancing the reliability of synthetic data auditing and the credibility of model evaluations.
Existing contamination detection methods rely on assumptions such as access to training data, handcrafted statistics, or predefined labels, limiting their applicability in real-world settings. This work proposes a novel approach that dispenses with such assumptions by constructing depth profiles via linear probes in residual streams, introducing a metric termed “excess separability,” and combining label permutation tests, item-wise bootstrapping, and size-matched placebo control sets to detect whether a model has been exposed to test data while controlling for confounding factors. The method achieves a substantially reduced false positive rate—down to 0.02—rejects fragile ablations, and demonstrates effectiveness on real Transformer models: it finds no evidence of contamination across four Pile subsets. All implementation and auditing code is publicly released.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.