workload isolation

Designs, implements, and evaluates mechanisms and formal or empirical models that establish and enforce boundaries between co‑located workloads to prevent performance interference, unintended side effects, or unauthorized information flow. This work covers resource partitioning and admission/scheduling policies, sandboxing and access-control techniques, architectural or runtime enforcement, and methods to specify, measure, and verify the strength of isolation across CPU, memory, storage, network, and software/hardware layers.

workloadisolation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.81
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$210K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Locked In, Leaked Out: Measuring Isolation via Kernel Locks

Jul 28, 2025
A
Anjali
🏛️ University of Wisconsin-Madison

In multi-tenant environments, lock contention on shared kernel data structures—such as filesystem journals and page allocators—is a primary cause of performance interference and isolation failure. This paper proposes the first method to quantitatively assess software-layer isolation by measuring kernel lock contention: leveraging system-level monitoring to precisely track acquisition latency and blocking frequency of diverse kernel locks under concurrent multi-workload execution, thereby establishing a fine-grained isolation analysis framework. Experiments demonstrate that the method accurately identifies isolation bottlenecks across kernel subsystems in virtual machines and containers, revealing filesystem logging and memory management as the dominant sources of interference. Unlike conventional black-box performance profiling, this work pioneers modeling lock-level synchronization behavior as a principled metric for isolation—providing an interpretable, reproducible, and low-level analytical foundation for designing, optimizing, and scheduling resources in multi-tenant systems.

Assess isolation in shared kernel data structuresIdentify common sources of workload interferenceMeasure performance interference via kernel locks

Cloud benchmarking suffers from performance variability due to multi-tenant resource contention. While the existing Duet approach mitigates cross-VM interference via co-locating two benchmark versions on the same VM, it overlooks implicit intra-VM workload competition. This work systematically evaluates the noise robustness of multiple isolation mechanisms—cgroups/CPU pinning, Docker containers, and Firecracker microVMs—under synchronous benchmarking, using controlled noise injection to emulate resource contention. Results show that Docker containers, owing to shared kernel resources and inherent cgroups scheduler limitations, significantly amplify performance variance in synchronous settings, rendering them unsuitable for high-precision benchmarking. In contrast, process-level isolation and microVMs effectively suppress interference, reduce false positives, and improve measurement consistency. The study uncovers latent performance risks of containerized environments in synchronous benchmarking and provides empirical evidence and practical guidance for selecting isolation strategies in cloud-native benchmarking.

Comparing cgroups, Docker containers and MicroVMs against noise interferenceEvaluating isolation mechanisms for synchronized cloud benchmarking workloadsMitigating performance variability from multi-tenant resource contention

This study addresses a critical side-channel vulnerability in cloud environments where containers and virtual machines, despite employing software-based isolation mechanisms, remain susceptible to cross-tenant information leakage through the shared host page cache. The authors systematically evaluate the page cache risks across diverse runtime environments—including Docker, gVisor, Kata, and QEMU/KVM—under shared storage conditions, leveraging unprivileged timing measurements to infer cache residency across isolation boundaries. Their work is the first to demonstrate the pervasive nature of page cache leakage in modern isolation architectures and integrates this attack vector into a unified framework for OS-mediated microarchitectural timing side channels. Experiments confirm that the attack succeeds whenever the I/O path involves shared cacheable file objects, while mitigation strategies such as direct I/O or dedicated block devices significantly suppress the signal. The approach successfully recovers coarse-grained activity patterns from a real-world WordPress+MySQL deployment.

cloud isolationisolation failurepage-cache

This work addresses the security risks arising from the lack of process-level isolation in CXL-based shared memory systems. To bridge the gap between host-level and process-level isolation, the authors propose a hardware-software co-design that, for the first time, enables fine-grained process-level access control and execution context authentication in disaggregated CXL memory. The solution integrates hardware-based attestation, memory access control, and a lightweight caching acceleration mechanism. Evaluated using an SST-based simulation model, the system supports up to 127 concurrent processes with only 3.3% performance overhead, significantly enhancing both the security and efficiency of shared memory architectures.

CXLmemory disaggregationprocess-level isolation

This work addresses the programming challenges and error-proneness introduced by Arm’s POE2 architecture, which employs a complex spatiotemporal permission mechanism yet lacks a unified security model. We propose the first general-purpose secure programming model tailored for POE2, abstracting away the intricacies of its spatial and temporal indexing and encapsulating hardware features such as memory protection keys, dedicated registers, and table structures. By doing so, our model significantly simplifies permission management while preserving POE2’s strong security guarantees. It naturally supports common intra-process isolation patterns used in software partitioning, enabling developers to construct secure isolated systems more efficiently and with fewer errors.

architectural complexityintra-process isolationmemory protection keys

Latest Papers

What's happening recently
View more

Current research on the security of LLM-agent systems remains fragmented, lacking a unified framework to explain the common root causes and propagation mechanisms underlying failures such as prompt injection and tool misuse. This work establishes *isolation* as a first-class principle for system security and introduces a boundary-centric taxonomy comprising five boundary types: user–agent, agent–tool, agent–execution, agent–agent, and system–environment. By systematically modeling failure pathways and defense strategies through structured review and cross-domain analysis, the study reveals that security failures predominantly originate from insufficient isolation and follow distinct cross-boundary attack propagation patterns. The paper thus provides a cohesive theoretical foundation and a construction-oriented research agenda centered on isolation for designing highly secure agent systems.

boundary failureisolationLLM-agent system safety

This work proposes a secure and transparent egress solution for sandboxed workloads to ensure multi-tenant isolation and fair resource allocation. By constructing a defense-in-depth architecture that integrates eBPF-based packet filtering, GENEVE overlay networking, and distributed egress proxies, the system enables policy-driven network access control with low overhead. A novel two-tier policy enforcement mechanism is introduced: at the lower layer, eBPF enforces precise bandwidth rate-limiting using the Earliest Departure Time (EDT) algorithm, while the upper layer provides protections against connection exhaustion and port depletion. Deployed across all Snowflake regions, the system supports petabyte-scale data transfers and low-latency external integrations, meeting the stringent security and performance requirements of large-scale production environments.

external connectivitymulti-tenant isolationresource fairness

This study addresses the security risks—such as privilege expiration, unauthorized escalation, and capability reuse—introduced by untrusted hosts in resource disaggregation architectures. To mitigate these threats, this work proposes a SmartNIC-based distributed capability access control system. The system introduces a novel capability mechanism that integrates host-independent verification, authoritative execution at the resource side, and efficient revocation. By leveraging FPGA hardware isolation, it enables independent privilege validation and secure revocation across compute and resource nodes, supported by a formal security proof. Experimental evaluation of the prototype demonstrates a throughput of 89.5 Gbit/s, with the revocation of a 128-node capability subtree requiring only 528 nanoseconds. These results confirm that the proposed architecture successfully unifies high performance with robust security isolation for disaggregated systems.

authority revocationcapability-based access controlresource disaggregation

Hot Scholars

MC

Matteo Cederle

PhD Student, University of Padova
reinforcement learningartificial intelligencesmart mobilitydeep learning
YW

Yite Wang

Research Scientist, Snowflake
Large Language ModelEfficient Deep LearningComputer VisionNatural Language Processing
PP

Peter Pietzuch

Professor of Distributed Systems, Imperial College London
Distributed SystemsSystemsData Management
NL

Noura Limam

University of Waterloo
Network operations and managementCommunication and Networking in the CloudNetwork Performance
ZZ

Zhenxing Zhang

School of computing, Dublin City University
machine learningcomputer visioninformation retrieval