🤖 AI Summary
This work addresses the challenge of transparently scaling multithreaded applications on existing CXL-based distributed shared memory (DSM) systems, which are hindered by the need for manual code modifications, static data placement policies, and sub-microsecond page fault overheads. To overcome these limitations, the authors propose a co-designed operating system and runtime that establish a global, unified address space, enabling transparent application scaling without source-code changes. Key innovations include a dynamic, latency-driven data placement strategy and an elastic page management mechanism that supports online page merging and splitting while exploiting spatial locality. Experimental evaluation across 15 configurations and diverse workloads demonstrates performance improvements of 1.5–2.2× over a pure CXL baseline and 1.1–2.2× over state-of-the-art hybrid DSM systems, achieving near-linear scalability.
📝 Abstract
While CXL presents a promising hardware substrate for Distributed Shared Memory (DSM), seamlessly scaling multithreaded applications across multiple nodes remains a formidable challenge. Existing CXL-based DSMs fall short: they require manual code modifications to share non-heap data, employ rigid data placement policies that fail under diverse and dynamic workloads, and suffer from severe page-fault processing overheads in sub-microsecond ($μ\mathrm{s}$) environments.
We present xDSM, a full-space, elastic DSM system built over CXL that transparently scales unmodified multithreaded applications. To eliminate the burden of manual code rewrites, xDSM employs an OS-runtime co-design that establishes a globally coordinated address space, seamlessly sharing all memory segments. To mask CXL access penalties, xDSM abandons static placement rules in favor of a dynamic, latency-driven policy that actively balances data between local DRAM and CXL memory. Finally, to resolve the fundamental tension between high base-page fault overheads and severe huge-page false sharing, xDSM introduces spatial locality-aware elasticity, dynamically coalescing and splitting pages on the fly to amortize processing costs.
Evaluated across diverse workloads using 15 system configurations, xDSM outperforms CXL-only baselines by 1.5$\times$ to 2.2$\times$ and state-of-the-art hybrid DSMs by 1.1$\times$ to 2.2$\times$, while achieving near-linear scalability.