Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions

📅 2024-10-24
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
LLM evaluation is frequently compromised by train-test data contamination, yet existing contamination detection methods rely on unvalidated underlying assumptions. This paper systematically reviews 50 studies to propose the first taxonomy of assumptions in contamination detection, identifying eight core assumption categories and empirically evaluating three representative ones. We introduce a novel hypothesis-driven evaluation paradigm integrating membership inference attacks (MIAs), distributional shift analysis, and large-scale contamination simulation. Our findings reveal that state-of-the-art MIAs perform near-randomly on real LLM pretraining data; distributional shifts substantially degrade detection reliability; and LLMs preferentially learn statistical patterns over memorizing individual instances. These results challenge the default “instance memorization” assumption, offering both theoretical foundations and methodological guidance for trustworthy LLM evaluation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Search and Optimization: Evaluation and Analysis

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models (LLMs) have demonstrated great performance across various benchmarks, showing potential as general-purpose task solvers. However, as LLMs are typically trained on vast amounts of data, a significant concern in their evaluation is data contamination, where overlap between training data and evaluation datasets inflates performance assessments. Multiple approaches have been developed to identify data contamination. These approaches rely on specific assumptions that may not hold universally across different settings. To bridge this gap, we systematically review 50 papers on data contamination detection, categorize the underlying assumptions, and assess whether they have been rigorously validated. We identify and analyze eight categories of assumptions and test three of them as case studies. Our case studies focus on detecting direct, instance-level data contamination, which is also referred to as Membership Inference Attacks (MIA). Our analysis reveals that MIA approaches based on these three assumptions can have similar performance to random guessing, on datasets used in LLM pretraining, suggesting that current LLMs might learn data distributions rather than memorizing individual instances. Meanwhile, MIA can easily fail when there are data distribution shifts between the seen and unseen instances.
Problem

Research questions and friction points this paper is trying to address.

Assessing data contamination detection in LLMs.
Evaluating assumptions in data contamination detection methods.
Testing Membership Inference Attacks on LLM pretraining datasets.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematically review data contamination detection methods
Categorize and validate eight detection assumptions
Test Membership Inference Attacks on LLM datasets
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Washington | George Mason University
Y
Yujuan Fu
University of Washington
Ö
Özlem Uzuner
George Mason University
M
Meliha Yetisgen-Yildiz
University of Washington
F
Fei Xia
University of Washington