Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

📅 2026-08-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical flaw in existing security evaluations, wherein stored labels are erroneously treated as ground-truth behavioral facts, leading to processing leaks and construct validity threats. To rectify this, the authors propose a seven-stage integrity chain and an endpoint integrity checker that jointly audit historical label assignment mechanisms through execution trace tracing, blind processing reconstruction, double-blind consistency review, and semantic boundary analysis. This approach exposes the non-terminal nature of labels and reconstructs the measurement foundation of security assessments. Empirically, the method corrects 58 mislabeled instances, eliminates all false-positive attack records in the v2 census while preserving genuine privilege escalation cases, and achieves consensus across reviewers on 96 explainable requests.
📝 Abstract
Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli. Two schema treatments were delivered, but the planned external payload-family corpus was not. The historical grader exhibited direct treatment leakage: treatment metadata gated the ATTACK_SUCCESS class, so fixed behavior could change class under treatment relabeling. A treatment-blind reconstruction corrects 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to authorized benign completions while preserving three verified protected-data transfers and one separate unauthorized-forwarding case. The locked v2 census contains exactly zero ATTACK_SUCCESS records, while the forwarding case remains a HIJACK_ATTEMPT at a semantic boundary concerning objective completion. A dual-reviewer blinded concordance review of all 96 requests deemed structurally interpretable by locked v2 produced identical reviewer-consensus classes but differed from the locked codebook on four construct-boundary cases. We contribute a seven-link Integrity Chain and an executable, scope-bounded endpoint-integrity linter. The result is a campaign-bounded measurement audit, not a population attack-rate, model-ranking, defense-efficacy, or causal estimate.
Problem

Research questions and friction points this paper is trying to address.

treatment leakage
construct validity
security evaluation
labeling bias
agent behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

treatment leakage
construct validity
endpoint integrity
label reconstruction
integrity chain
R
Rana Muhammad Ahmed
Department of Computer Science, Bahria University, Islamabad, Pakistan
S
Sabahat Abbas
Department of Computer Science, Bahria University, Islamabad, Pakistan