PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of existing agent skill scanners to detect concealed bytecode cache attacks due to the disconnect between inspection and execution. It reveals Python cache replacement vulnerabilities and introduces Execution-Aware Verification (EAV), a novel mechanism that extends static analysis to dynamically compiled artifacts. By constructing typed execution graphs, performing grounded behavioral analysis, and leveraging trusted reproduction techniques, EAV enables deep validation of runtime artifacts, effectively bridging the semantic gap between inspection and execution. Experimental results demonstrate that the proposed approach successfully detects all 100 attack samples, achieving a recall rate of 92.8% and a false positive rate of 10.0% across five distinct attack families.
📝 Abstract
Agent skills combine instructions with executable resources, giving third-party packages access to an agent's runtime. Existing skill scanners inspect documentation and visible source, but Python may execute a bundled bytecode cache with different behavior. We study this gap between inspection and execution through PyCache Trap, which pairs benign source with a substituted cache accepted by the loader and connects it to a task-relevant invocation. Scanner-guided rewriting changes the invocation wording while preserving the cache body, separating package admission from recognition of the concealed behavior. Across 100 skills and seven scanners, PyCache Trap achieves 94-100% attack success, with no semantic recognition of the cache-resident behavior. We propose execution-aware validation (EAV) to connect inspected instructions, scripts, imports, and runtime artifacts in a typed execution graph. EAV combines grounded behavioral analysis with trusted reproduction of compiled artifacts. It detects all 100 evaluated source-present cache substitutions and reaches 92.8% Recall at 10.0% FPR across five attack families and 200 benign skills. The results support checking the executable artifacts a runtime can select as part of skill admission, within the supported loaders and code-object normalization. The code is released at https://github.com/leo0481/PyCacheTrap.
Problem

Research questions and friction points this paper is trying to address.

Agent Skill Security
Inspection-Execution Gap
PyCache Trap
Bytecode Cache Substitution
Skill Scanner Evasion
Innovation

Methods, ideas, or system contributions that make the work stand out.

PyCache Trap
Execution-Aware Validation
Agent Skill Scanners
Bytecode Cache Substitution
Inspection-Execution Gap
🔎 Similar Papers
No similar papers found.