in-storage graph processing

Designs and implements computation kernels, runtimes, and scheduling policies that execute graph algorithms directly on storage devices (e.g., flash/SSD controllers), including lightweight in-device kernels, batching and memory-aware scheduling to fit limited device memory, and mechanisms to offload low-reuse data accesses from the host to storage. Analyzes performance and device-level constraints (latency, throughput, wear, durability, consistency) and integration with host I/O stacks and data layouts to maximize throughput and minimize host-side data movement.

in-storagegraphprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the memory wall between processors and storage, where data movement has become a critical performance bottleneck. Existing computational storage solutions struggle to scale due to programming complexity, ecosystem fragmentation, and thermal/power constraints. To overcome these limitations, the authors propose a reversible computational storage architecture that enables dynamic migration of WebAssembly-compiled storage executables between the host and CXL SSDs. The design leverages CXL.mem’s cache coherence for seamless state sharing and introduces a zero-copy drain-and-switch protocol to manage thermal and power constraints. An agility-aware scheduler elastically dispatches compute tasks based on runtime conditions. Evaluations on both FPGA prototypes and commercial computational storage devices demonstrate up to 2× higher throughput and 3.75× lower write latency without requiring application modifications, effectively transforming rigid thermal limits into tunable performance trade-offs.

computational storageCXL SSDsdata movement bottleneck

Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems

Mar 13, 2025
FK
Fabian Knorr
🏛️ University of Innsbruck

SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.

Enhancing memory allocation and concurrency in SYCL programs on accelerator clusters.Optimizing scheduling for high-level parallel programs on multi-GPU systems.Reducing delays in distributed-memory applications through graph-based representations.

Revisiting the Design of In-Memory Dynamic Graph Storage

Feb 10, 2025
JS
Jixian Su
🏛️ Shanghai Jiao Tong University | Huawei Cloud | National University of Singapore

This paper addresses performance bottlenecks of in-memory dynamic graph storage (DGS) under high-concurrency read/write workloads, identifying three fundamental issues: (1) excessive memory redundancy—3.3× to 8.9× higher than CSR; (2) poor cache efficiency due to neglect of modern memory access patterns; and (3) severe contention and versioning overhead on high-degree vertices caused by fine-grained concurrency control. To systematically analyze these challenges, the authors propose a unified DGS abstraction model and a configurable, multi-dimensional benchmarking framework—the first to quantitatively evaluate trade-offs among throughput, latency, memory footprint, and cache behavior across mainstream DGS designs. Empirical results demonstrate that fine-grained versioning is fundamentally unsuitable for highly concurrent dynamic graph workloads. The study provides both theoretical foundations and practical guidance for designing next-generation DGS architectures that are low-overhead, cache-friendly, and scalable.

Evaluates in-memory dynamic graph storage effectivenessHighlights space overhead and concurrency control issuesIdentifies performance factors in graph storage methods

This study addresses the high software engineering overhead and severe write amplification inherent in conventional data placement mechanisms for NVMe SSDs. To this end, we propose a cross-layer explicit data placement mechanism based on the Reclaim Units abstraction. By leveraging NVMe Flexible Data Placement (FDP) technology to construct a Linux block I/O-compatible interface, our approach enables low-intrusion mapping for MySQL and RocksDB, optimizing storage performance without imposing sequential write constraints or requiring host-side garbage collection. Experimental results demonstrate that the proposed scheme significantly reduces the end-to-end write amplification factor under heavy workloads while improving both quality of service and throughput. Furthermore, it supports seamless integration within the open-source ecosystem, exhibiting substantial potential for rapid industrial deployment.

Data PlacementFlexible Data Placement (FDP)NVMe SSD

This paper addresses frequent cold starts and low resource utilization in edge-based Serverless computing, caused by resource constraints and hardware heterogeneity. To tackle these challenges, we propose KiSS—a static, container-size-aware memory management strategy. Its core contribution is the first static memory partitioning mechanism that jointly models container size and invocation frequency: the memory pool is partitioned into two isolated regions—“small & high-frequency” and “large-resource”—enabling interference-aware resource isolation. We validate KiSS through discrete-event simulation and real-world edge deployments. Under typical workloads, KiSS reduces cold-start occurrences by 60% and function drop rates by 56.5%, significantly improving responsiveness and stability of edge Serverless systems.

Minimize inter-function interference for efficient resource utilizationOptimize memory management in edge-cloud continuumReduce cold-start in serverless edge environments

Latest Papers

What's happening recently
View more

Traditional graph processing systems are constrained by monolithic architectures, where tightly coupled resources lead to low utilization. Existing memory-disaggregated approaches suffer from poor scalability and high cache overhead. This work proposes DMG, the first practical memory-disaggregated graph processing system, which introduces a disaggregation-friendly graph storage layout, an adaptive update coordination mechanism, and a two-level load management strategy to enable efficient graph access, low-overhead update propagation, and dynamic load balancing. DMG is the first system to support elastic scaling across multiple compute and memory nodes while significantly reducing cache requirements without sacrificing performance. Experimental results demonstrate that DMG achieves up to 4.9× higher performance and reduces cache footprint by up to 18.9× compared to the state-of-the-art systems.

cache efficiencygraph processingmemory-disaggregated

This study addresses the challenge that static compilation and limited memory bandwidth on mobile NPUs render conventional KV caching impractical for on-device large language model (LLM) inference. To overcome this, we propose a compute-storage co-design framework for KV cache reuse. Specifically, the method incorporates a selective recomputation mechanism within static graphs, a dynamic programming-based cross-graph scheduler, and a hierarchical KV manager integrating tree hashing with semantic matching. Additionally, a two-dimensional pipelining technique is employed to hide data transfer latency. Experimental results demonstrate that, compared with non-reuse and prefix-only caching baselines, the proposed framework reduces time-to-first-token (TTFT) by 40%–60%, substantially improving LLM inference efficiency on mobile devices.

KV cache reuseMemory bandwidth constraintMobile NPU

This work addresses the performance degradation in large language model (LLM) inference on multi-chiplet NUMA GPUs, where non-uniform memory access and inter-chiplet communication induce kernel latency and poor memory locality. The study presents the first systematic classification of operand sharing patterns across workgroups in LLM kernels—categorized as global, partial, or private—and leverages memory trace analysis, workgroup-level access modeling, and cycle-accurate simulation to demonstrate that each pattern necessitates distinct data placement and scheduling strategies. Building on these insights, the authors propose a sub-group-aware co-scheduling mechanism coupled with an optimized data layout scheme, which significantly enhances execution efficiency and memory locality for LLM kernels on multi-chiplet GPU architectures.

kernel latencyLLMmemory locality

Hot Scholars

MS

Mohammad Sadrosadati

Senior Researcher and Lecturer, ETH Zürich
Heterogeneous ComputingProcessing-In-MemoryMemory SystemsInterconnection Networks
YX

Yantuan Xian

Kunming University of Science and Technology
machine learningnatural language processingtext mining
HM

Harun Mustafa

Johns Hopkins University
Computational BiologyMetagenomicsAlgorithms