🤖 AI Summary
This study addresses the fragmentation and limited comparability of existing evaluation criteria for extraction attacks and defenses against large language models (LLMs) by constructing the first standardized benchmark spanning the entire lifecycle. The proposed framework encompasses six categories of black-box attacks, ten defense mechanisms, and adaptive rewriting attacks. Through unified configuration management and query budget control, it enables joint multidimensional measurement of capability, fidelity, and quality. This project establishes a highly reproducible foundation for comparative analysis, systematically revealing the limitations of current defenses under adaptive attacks. Ultimately, this work formalizes a standardized evaluation paradigm for LLM security research, facilitating rigorous and consistent assessment across diverse threat models and mitigation strategies.
📝 Abstract
Large language models (LLMs) deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access assumptions, model configurations, query budgets, and security objectives, limiting comparability across methods. To address this gap, we introduce a unified benchmark covering six extraction attacks, ten defenses, and two adaptive attacks that paraphrase or back-translate protected responses before surrogate training. The benchmark controls model configurations, query data, budgets, and held-out evaluation conditions within each comparison while preserving attack-specific querying and training procedures. We measure surrogate capability, fidelity to the victim, output quality using Rep-4, and query-budget sensitivity; defenses use their own security metrics paired with surrogate performance. For the adaptive attacks, we jointly measure provenance-detector scores and the capability and fidelity of surrogates trained on rewritten responses. The benchmark thus provides a reproducible basis for comparing extraction methods and their interactions with defenses under text-only access. Code and artifacts are available at https://github.com/sliu11-byte/MEA-Bench.