🤖 AI Summary
Arm CCA currently supports only VM-level isolation, rendering it ineffective against intra-process vulnerabilities such as Heartbleed; existing fine-grained isolation schemes suffer from high performance overhead or incompatibility with Confidential Virtual Machines (CVMs). This paper proposes a three-tiered region isolation model that enables infinitely lightweight, intra-process security domains, achieving fine-grained memory protection and coordinated defense across user- and kernel-space. We introduce the first CCA extension architecture supporting dynamic domain creation, integrated with a lightweight user-space Code Pointer Integrity (CPI) mechanism—ensuring strong security guarantees while preserving full CVM compatibility. Evaluated on both Arm simulator and physical development boards, our prototype incurs only ~20% average performance overhead while sustaining 95% throughput. It effectively mitigates Heartbleed, session key leakage, and sensitive data exfiltration from KV stores and non-volatile memory.
📝 Abstract
Arm Confidential Computing Architecture (CCA) currently isolates at the granularity of an entire Confidential Virtual Machine (CVM), leaving intra-VM bugs such as Heartbleed unmitigated. The state-of-the-art narrows this to the process level, yet still cannot stop attacks that pivot within the same process, and prior intra-enclave schemes are either too slow or incompatible with CVM-style isolation. We extend CCA with a three-tier zone model that spawns an unlimited number of lightweight isolation domains inside a single process, while shielding them from kernel-space adversaries. To block domain-switch abuse, we also add a fast user-level Code-Pointer Integrity (CPI) mechanism. We developed two prototypes: a functional version on Arm's official simulator to validate resistance against intra-process and kernel-space adversaries, and a performance variant on Arm development boards evaluated for session-key isolation within server applications, in-memory key-value protection, and non-volatile-memory data isolation. NanoZone incurs roughly a 20% performance overhead while retaining 95% throughput compared to the system without fine-grained isolation.