TDDBench: A Benchmark for Training data detection

📅 2024-11-05
🏛️ International Conference on Learning Representations
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
A unified, multimodal benchmark for training data detection (TDD) evaluation is currently lacking. Method: This paper introduces TDDBench—the first cross-modal TDD benchmark—comprising 13 datasets across three modalities (images, tabular data, and text), and systematically evaluates 21 TDD methods under four paradigms: black-box, white-box, gradient-based, and reconstruction-based. Contribution/Results: TDDBench establishes the first standardized, multidimensional evaluation framework, jointly measuring accuracy (accuracy/F1-score), efficiency (inference latency), and resource cost (memory footprint). Experiments reveal that existing TDD methods exhibit limited detection capability, particularly under cross-modal settings and realistic noise conditions. The benchmark platform is fully open-source, modular, and extensible, enabling reproducible algorithm comparison, quantitative analysis, and practical deployment of TDD techniques.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal VisionData Mining & Knowledge Management: Mining of Visual, Multimedia & Multimodal Data

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
Training Data Detection (TDD) is a task aimed at determining whether a specific data instance is used to train a machine learning model. In the computer security literature, TDD is also referred to as Membership Inference Attack (MIA). Given its potential to assess the risks of training data breaches, ensure copyright authentication, and verify model unlearning, TDD has garnered significant attention in recent years, leading to the development of numerous methods. Despite these advancements, there is no comprehensive benchmark to thoroughly evaluate the effectiveness of TDD methods. In this work, we introduce TDDBench, which consists of 13 datasets spanning three data modalities: image, tabular, and text. We benchmark 21 different TDD methods across four detection paradigms and evaluate their performance from five perspectives: average detection performance, best detection performance, memory consumption, and computational efficiency in both time and memory. With TDDBench, researchers can identify bottlenecks and areas for improvement in TDD algorithms, while practitioners can make informed trade-offs between effectiveness and efficiency when selecting TDD algorithms for specific use cases. Our large-scale benchmarking also reveals the generally unsatisfactory performance of TDD algorithms across different datasets. To enhance accessibility and reproducibility, we open-source TDDBench for the research community.
Problem

Research questions and friction points this paper is trying to address.

Lack of comprehensive benchmark for Training Data Detection (TDD) methods evaluation
Need to assess TDD effectiveness across multiple data modalities and paradigms
Unsatisfactory performance of current TDD algorithms across diverse datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces TDDBench for TDD evaluation
Benchmarks 21 methods across four paradigms
Open-sources TDDBench for research community
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Science and Technology of China | The Hong Kong University of Science and Technology
Zhihao Zhu
Zhihao Zhu
University of Science and Technology of China
Machine Learning PrivacyRecommender SystemGraph Neural Network
Y
Yi Yang
The Hong Kong University of Science and Technology
D
Defu Lian
University of Science and Technology of China