A Survey on Data Contamination for Large Language Models

📅 2025-02-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses data contamination in large language model (LLM) evaluation—unintended overlap between training data and evaluation benchmarks that inflates performance metrics and misleads generalization assessment. Methodologically, it introduces the first taxonomy of contamination detection based on model information dependency: white-box (leveraging gradients or memory traces), gray-box (assessing output consistency), and black-box approaches. It further proposes a novel paradigm of dynamic benchmark construction and LLM-driven contamination-free evaluation, complemented by an integrated mitigation strategy comprising data updating, rewriting, and prevention. The contributions are threefold: (1) exposing systemic vulnerabilities in current LLM evaluation practices; (2) establishing the first comprehensive, classification-based governance framework spanning contamination detection, unbiased evaluation, and proactive prevention; and (3) providing both theoretical foundations and practical guidelines for developing more rigorous and trustworthy LLM evaluation protocols.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsSearch and Optimization: Evaluation and Analysis

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Recent advancements in Large Language Models (LLMs) have demonstrated significant progress in various areas, such as text generation and code synthesis. However, the reliability of performance evaluation has come under scrutiny due to data contamination-the unintended overlap between training and test datasets. This overlap has the potential to artificially inflate model performance, as LLMs are typically trained on extensive datasets scraped from publicly available sources. These datasets often inadvertently overlap with the benchmarks used for evaluation, leading to an overestimation of the models' true generalization capabilities. In this paper, we first examine the definition and impacts of data contamination. Secondly, we review methods for contamination-free evaluation, focusing on three strategies: data updating-based methods, data rewriting-based methods, and prevention-based methods. Specifically, we highlight dynamic benchmarks and LLM-driven evaluation methods. Finally, we categorize contamination detecting methods based on model information dependency: white-Box, gray-Box, and black-Box detection approaches. Our survey highlights the requirements for more rigorous evaluation protocols and proposes future directions for addressing data contamination challenges.
Problem

Research questions and friction points this paper is trying to address.

Data contamination in LLM evaluation
Methods for contamination-free evaluation
Detection approaches for data contamination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic benchmarks for contamination-free evaluation
LLM-driven methods for accurate assessment
White-Box, Gray-Box, Black-Box detection approaches
Y
Yuxing Cheng
College of Software, Jilin University
Y
Yi Chang
School of Artificial Intelligence, Jilin University, Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China, International Center of Future Science, Jilin University
Y
Yuan Wu
School of Artificial Intelligence, Jilin University