🤖 AI Summary
This study addresses the lack of systematic, engine-level security comparisons among existing AI code sandboxing solutions, which hinders reliable assessment of their isolation capabilities and associated risks. The work proposes a novel cross-dimensional analytical framework that evaluates five mainstream sandbox products across six criteria: host attack surface, information leakage potential, stackability of defense-in-depth mechanisms, historical CVE records, patching cadence, and upstream fuzz testing maturity. Through comprehensive attack surface mapping, information flow analysis, CVE mining, patch latency tracking, and fuzzing ecosystem evaluation, the research uncovers clear security boundaries between distinct engine architectures and significant disparities in security practices among functionally similar products. Notably, patch delays range from zero to over 471 days, and the optimal security configuration—combining microVM isolation with continuous public fuzz testing—remains unadopted. To guide practical deployment, the authors introduce a threat-model-adaptive selection matrix as a nuanced alternative to simplistic rankings.
📝 Abstract
This paper reads six engine-level measurements together -- 1.1 host attack surface, 1.2 information leakage, 1.3 defense-in-depth stackability, 1.4 public CVE history, 1.5 patch cadence, and 1.6 upstream fuzzing posture -- to describe how five AI-sandbox products isolate guest code from the host kernel. No single axis is a sufficient basis for a comparative judgement; the cross-axis reading is the load-bearing analysis.
Three high-level findings: (1) engine classes (microVM, userspace kernel, OCI container) separate cleanly on every architectural axis, but products within a class do not; (2) product pin policy is the dominant operator-facing variable -- engine-side patch latency aggregates to ~0 days for coordinated disclosures, while downstream lag spans 0 days to 471+ days to "opaque" to infinity; (3) fuzzing investment splits into three tiers, and the strongest combination -- microVM x continuous public fuzzer -- is unoccupied in this set, leaving the "0 published CVEs x no upstream fuzzer x no academic study" intersection structurally unmeasured.
We report per-axis orderings, per-product portraits, and a threat-model qualification matrix; no overall ranking is proposed. Companion repository (code, Apache-2.0): https://github.com/orbitalab/RnD-ai-sandboxes-sec-study-part-1. License: CC BY 4.0.