Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragmentation and limited comparability of existing evaluation criteria for extraction attacks and defenses against large language models (LLMs) by constructing the first standardized benchmark spanning the entire lifecycle. The proposed framework encompasses six categories of black-box attacks, ten defense mechanisms, and adaptive rewriting attacks. Through unified configuration management and query budget control, it enables joint multidimensional measurement of capability, fidelity, and quality. This project establishes a highly reproducible foundation for comparative analysis, systematically revealing the limitations of current defenses under adaptive attacks. Ultimately, this work formalizes a standardized evaluation paradigm for LLM security research, facilitating rigorous and consistent assessment across diverse threat models and mitigation strategies.
📝 Abstract
Large language models (LLMs) deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access assumptions, model configurations, query budgets, and security objectives, limiting comparability across methods. To address this gap, we introduce a unified benchmark covering six extraction attacks, ten defenses, and two adaptive attacks that paraphrase or back-translate protected responses before surrogate training. The benchmark controls model configurations, query data, budgets, and held-out evaluation conditions within each comparison while preserving attack-specific querying and training procedures. We measure surrogate capability, fidelity to the victim, output quality using Rep-4, and query-budget sensitivity; defenses use their own security metrics paired with surrogate performance. For the adaptive attacks, we jointly measure provenance-detector scores and the capability and fidelity of surrogates trained on rewritten responses. The benchmark thus provides a reproducible basis for comparing extraction methods and their interactions with defenses under text-only access. Code and artifacts are available at https://github.com/sliu11-byte/MEA-Bench.
Problem

Research questions and friction points this paper is trying to address.

Model Extraction
Large Language Models
Black-Box Attacks
Defenses
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model Extraction
Large Language Models
Unified Benchmark
Adaptive Attacks
Defenses
💼 Related Jobs
No related jobs found.
S
Shuze Liu
Florida State University
K
Kaixiang Zhao
Brigham Young University
R
Runyang Xu
University of Michigan, Ann Arbor
J
Jingzhi Chen
State University of New York at Buffalo
N
Nathan Wu
Wake Forest University
Y
Yu Wang
University of Georgia
Yushun Dong
Yushun Dong
Assistant Professor, Department of Computer Science, Florida State University
AI SecurityAI IntegrityGraph Machine LearningLLMs