MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

📅 2026-04-30
📈 Citations: 0
Influential: 0
📄 PDF

career value

230K/year
🤖 AI Summary
This work addresses a non-adversarial information flow risk in multi-server Model Context Protocol (MCP) agents, where legitimate tool compositions may inadvertently propagate credentials across trust boundaries. The authors propose the first controllable benchmarking framework that combines honeypot-based taint tracking, coverage-oriented environment design, and Credential Redaction Stratification (CRS) to effectively distinguish between task-essential and policy-violating credential propagation paths. Evaluated across five models and 147 tasks, the study reveals violation rates ranging from 11.5% to 41.3%, predominantly through browser-mediated data flows. A novel prompting-based mitigation strategy reduces such violations by 97% while preserving 80.5% of functional utility.
📝 Abstract
Multi-server MCP agents create an information-flow control problem: faithful tool composition can turn individually benign read/write permissions into cross-boundary credential propagation -- a structural side effect of workflow topology, not necessarily malicious model behavior. We present MCPHunt, to our knowledge the first controlled benchmark that isolates non-adversarial, verbatim credential propagation across multi-server MCP trust boundaries, with three methodological contributions: (1) canary-based taint tracking that reduces propagation detection to objective string matching; (2) an environment-controlled coverage design with risky, benign, and hard-negative conditions that validates pipeline soundness and controls for credential-format confounds; (3) CRS stratification that disentangles task-mandated propagation (faithful execution of verbatim-transfer instructions) from policy-violating propagation (credentials included despite the option to redact). Across 3,615 main-benchmark traces from 5 models spanning 147 tasks and 9 mechanism families, policy-violating propagation rates reach 11.5--41.3% across all models. This propagation is pathway-specific (25x cross-mechanism range) and concentrated in browser-mediated data flows; hard-negative controls provide evidence that production-format credentials are not necessary -- prompt-directed cross-boundary data flow is sufficient. A prompt-mitigation study across 3 models reduces policy-violating propagation by up to 97% while preserving 80.5% utility, but effectiveness varies with instruction-following capability -- suggesting that prompt-level defenses alone may not suffice. Code, traces, and labeling pipeline are released under MIT and CC BY 4.0.
Problem

Research questions and friction points this paper is trying to address.

multi-server MCP agents
cross-boundary data propagation
credential leakage
information-flow control
trust boundaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

MCPHunt
cross-boundary data propagation
canary-based taint tracking
credential leakage
multi-server MCP agents
H
Haonan Li
1Key Laboratory of Intraplate Volcanoes and Earthquakes (China University of Geosciences, Beijing), Ministry of Education, Beijing 100083, China; 2School of Geophysics and Information Technology, China University of Geosciences, Beijing 100083, China
Tianjun Sun
Tianjun Sun
Rice University
personnel selectionindividual differencesresearch methodsapplied psychometrics
Y
Yongqing Wang
1Key Laboratory of Intraplate Volcanoes and Earthquakes (China University of Geosciences, Beijing), Ministry of Education, Beijing 100083, China; 2School of Geophysics and Information Technology, China University of Geosciences, Beijing 100083, China
Q
Qisheng Zhang
1Key Laboratory of Intraplate Volcanoes and Earthquakes (China University of Geosciences, Beijing), Ministry of Education, Beijing 100083, China; 2School of Geophysics and Information Technology, China University of Geosciences, Beijing 100083, China