COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to disentangle systemic safeguards from model vulnerabilities, precluding the isolated evaluation of large language models (LLMs) as Model Context Protocol (MCP) clients. This work proposes COPEX, a pioneering isolated testing framework that fixes the agent stack while exclusively substituting the underlying model. It encompasses 25 attacks across four entry points and introduces defense mechanisms such as hierarchical context injection and compositional input scanning. Experimental results demonstrate an average attack success rate of 64.4%, which is reduced by nearly half through combined defenses. This open-source benchmark precisely quantifies inherent model vulnerabilities, establishing a standardized paradigm for MCP security evaluation.
📝 Abstract
Large language models increasingly mediate tool use in Model Context Protocol (MCP) systems, where adversarial influence may enter through user instructions, tool schemas, tool outputs, or protocol messages. Existing benchmarks often evaluate deployed agents, conflating model susceptibility with guardrails, orchestration, and general task capability. We introduce COPEX (COntext Provider EXploitation), a controlled benchmark that isolates the model as an MCP client by fixing the surrounding agent stack and varying only the tool-selecting model. COPEX covers 25 attack types instantiated as 125 scenarios across four entry surfaces: model/agent, client, server/tool, and transport. Across nine models and 3,375 trials, the mean attack success rate is 64.4%, with surface-level means ranging from 58.3% to 71.4%. Some client- and transport-level attacks succeed partly outside the model's observation or control, separating system exposure from model susceptibility. Combined input and context scanning reduces mean attack success by 49.6% on an eight-attack defense subset relative to the undefended setting. The benchmark is available at https://github.com/inspire-center/copex.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Model Context Protocol
Adversarial Robustness
Benchmarking
Tool Use
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model Context Protocol
Adversarial Robustness
Benchmark
Large Language Models
Attack Surface
🔎 Similar Papers
N
Nahom Birhan
Old Dominion University
M
Mehrdad Rostamzadeh
Old Dominion University
S
Sidhant Narula
Old Dominion University
M
Mahmoud Nazzal
Old Dominion University
M
Mohammad Ghasemigol
Old Dominion University
Daniel Takabi
Daniel Takabi
Professor and Director of School of Cybersecurity, Old Dominion University
Trustworthy AIInformation Security & PrivacyUsable Security and Privacy