Score
Designs, implements, and evaluates mechanisms that confine code and data to restricted execution contexts—using process isolation, execution sandboxing, and trusted execution environments—and that separate privileges at the OS, hypervisor, or hardware level. Implements and analyzes resource isolation and timeboxing policies by reserving compute and network slots and enforcing CPU, memory, timing, and bandwidth limits to prevent interference, control information leakage, and bound runtime and resource consumption.
In multi-tenant environments, lock contention on shared kernel data structures—such as filesystem journals and page allocators—is a primary cause of performance interference and isolation failure. This paper proposes the first method to quantitatively assess software-layer isolation by measuring kernel lock contention: leveraging system-level monitoring to precisely track acquisition latency and blocking frequency of diverse kernel locks under concurrent multi-workload execution, thereby establishing a fine-grained isolation analysis framework. Experiments demonstrate that the method accurately identifies isolation bottlenecks across kernel subsystems in virtual machines and containers, revealing filesystem logging and memory management as the dominant sources of interference. Unlike conventional black-box performance profiling, this work pioneers modeling lock-level synchronization behavior as a principled metric for isolation—providing an interpretable, reproducible, and low-level analytical foundation for designing, optimizing, and scheduling resources in multi-tenant systems.
This work addresses the programming challenges and error-proneness introduced by Arm’s POE2 architecture, which employs a complex spatiotemporal permission mechanism yet lacks a unified security model. We propose the first general-purpose secure programming model tailored for POE2, abstracting away the intricacies of its spatial and temporal indexing and encapsulating hardware features such as memory protection keys, dedicated registers, and table structures. By doing so, our model significantly simplifies permission management while preserving POE2’s strong security guarantees. It naturally supports common intra-process isolation patterns used in software partitioning, enabling developers to construct secure isolated systems more efficiently and with fewer errors.
Traditional sandboxing mechanisms require application code refactoring, severely hindering deployment in legacy systems. This paper proposes Threadbox, a fine-grained, thread-level sandboxing framework that enables modular isolation and resource control for arbitrary functions without modifying application architecture. Its core innovation lies in lowering the sandbox boundary to the thread level, synergistically integrating runtime scheduling with OS-level resource isolation to deliver a lightweight, dynamic, and embeddable secure execution environment. Evaluation demonstrates that Threadbox effectively isolates sensitive operations with an average performance overhead of less than 8.2%. It significantly enhances sandbox flexibility, integrability, and practical applicability—advancing secure isolation toward modularity and runtime programmability.
Current cloud-sensitive workload protection relies on multi-layered nested isolation (e.g., confidential VMs, enclaves, sandboxes), leading to trusted computing base (TCB) bloat, complex end-to-end attestation, and cross-platform fragmentation. This work proposes Tyche, a unified isolation model introducing the first recursively nestable trust domain (TD) abstraction, managed by a lightweight security monitor that unifies resource isolation and remote attestation. Tyche operates atop commodity x86_64 and RISC-V hardware—requiring no specialized security extensions—and enables composable, verifiable, fine-grained isolation for confidential VMs, enclaves, and sandboxes. Its low-overhead, unified abstraction provides an SDK compatible with unmodified workloads, supporting demanding scenarios such as confidential inference at near-native Linux performance. We empirically validate Tyche’s cross-architecture portability and multi-tenant security guarantees.
This work addresses the security risks posed by AI agents frequently executing untrusted code on developer machines, where existing isolation mechanisms suffer from limitations in privilege requirements, performance overhead, and granularity of control. The authors propose a privilege-free, fine-grained process sandboxing architecture that compiles static security policies into kernel-enforced rules using Linux primitives such as seccomp and namespaces, while delegating dynamic decisions to a lightweight userspace supervisor. This approach enables rootless enforcement over filesystem, network, IPC, and system call access, supports time-of-check-to-time-of-use (TOCTOU)-safe validation and reversible file operations, and avoids dependencies on containers, cgroups, or images. Experimental results demonstrate a startup overhead of only ~5 ms, Redis performance matching bare-metal levels, and stage-based isolation of data, network, and untrusted content.
Current research on the security of execution environments for AI coding agents remains highly fragmented, lacking systematic integration and cross-disciplinary coordination. This work presents the first comprehensive survey of the field, analyzing 39 papers published between 2023 and 2026 and categorizing them into 17 thematic groups. Through CVE validation, cross-category comparison, and threat modeling, the study identifies critical disconnects among key areas such as isolation, access control, and time-of-check-to-time-of-use (TOCTOU) vulnerabilities, revealing five major research gaps. The analysis confirms four patched CVEs affecting production frameworks, quantifies the failure rate of existing mitigation strategies at 69%–98%, and uncovers that 17.1% of benign out-of-bound behaviors remain unaddressed by current mechanisms. Building on these findings, the paper proposes a unified research agenda to advance the field.
This study addresses a critical side-channel vulnerability in cloud environments where containers and virtual machines, despite employing software-based isolation mechanisms, remain susceptible to cross-tenant information leakage through the shared host page cache. The authors systematically evaluate the page cache risks across diverse runtime environments—including Docker, gVisor, Kata, and QEMU/KVM—under shared storage conditions, leveraging unprivileged timing measurements to infer cache residency across isolation boundaries. Their work is the first to demonstrate the pervasive nature of page cache leakage in modern isolation architectures and integrates this attack vector into a unified framework for OS-mediated microarchitectural timing side channels. Experiments confirm that the attack succeeds whenever the I/O path involves shared cacheable file objects, while mitigation strategies such as direct I/O or dedicated block devices significantly suppress the signal. The approach successfully recovers coarse-grained activity patterns from a real-world WordPress+MySQL deployment.
This work addresses the absence of a unified, verifiable runtime safety mechanism in existing MCP-style agents, where security decisions are fragmented across multiple components. To bridge this gap, the paper introduces HCP (Handle-Capability Protocol), a runtime framework that, while fully compatible with MCP workflows, formally defines eight execution-layer safety invariants for the first time. HCP enforces these invariants through a fine-grained access control model grounded in subjects, resources, capabilities, handles, and policies, explicitly ensuring critical properties such as subject binding, capability scoping, and data-flow authorization. Empirical evaluation demonstrates that HCP successfully blocks all attacks across ten benchmark scenarios while preserving auditable evidence, substantially outperforming baseline approaches. Microbenchmark results further indicate that policy operations incur an average latency of less than one millisecond.