Behind a Simple Read: Understanding Work and Waiting in the Linux I/O Stack

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the non-uniform performance behavior of Linux's simple read interface, where complex interactions among workload characteristics, processing time, and critical-path waiting impede optimization. To tackle this, the work proposes an end-to-end decomposition of the buffered read path, constructing a cross-layer wait-chain model alongside a "first-completion, subsequent-capacity" abstraction. These models are quantitatively characterized through kernel prototyping, MQSim simulation, and empirical SSD measurements. The proposed approach achieves block-layer prediction errors below 6.9%, reveals opportunities for out-of-order copying and parallel submission, and establishes a closed-loop framework spanning observational modeling to control optimization. Ultimately, this research provides quantifiable theoretical foundations and practical guidelines for optimizing layered I/O stacks.
📝 Abstract
A simple read interface provides uniform functional semantics, but not equally simple or predictable performance behavior. We use synchronous large-buffer reads in Linux as an observation window and decompose the buffered-read path end to end, distinguishing work volume, processing time, and critical-path exposure. We find that Linux reduces metadata work through large folios and overlaps most cold-read copying with device waiting, but these mechanisms depend on access advice, folio granularity, cache state, and backend execution. Using an experimental kernel prototype, we validate additional opportunities from out-of-order early copying and opportunistic parallelism, while showing that faster request submission does not improve end-to-end performance when SSD supply is already sufficient. To explain device waiting, we abstract first-completion wait $F$ and subsequent completion capacity $B_{CQ}$ from finite request-batch completion timelines. Measurements on a real SSD characterize how they vary with request size and batch size, while MQSim experiments connect them to internal device mechanisms and distinguish finite-batch completion from sustained throughput. We then construct a cross-layer wait chain along the actual folio--bio--request mapping. Independently calibrated device parameters predict waiting at the block layer and at read_pages with errors no greater than approximately 6.9% and 3.2%, respectively. Finally, four use cases apply the model and wait chain to parallel submission, Linux readahead, dependent-read layout, and polling versus sleeping. Their gains, no gains, and gain reversals form a closed loop from observation and modeling to explanation and control, providing a measurable basis for performance decisions in layered I/O stacks.
Problem

Research questions and friction points this paper is trying to address.

Linux I/O stack
buffered read performance
device waiting
cross-layer analysis
performance predictability
Innovation

Methods, ideas, or system contributions that make the work stand out.

I/O stack decomposition
cross-layer wait chain
device waiting model
out-of-order early copying
opportunistic parallelism
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yang Shen
College of Computer Science and Technology, National University of Defense Technology
Kai Lu
Kai Lu
Postdoc, Huazhong University of Science and Technology
Distributed storage systemskey-value storageAI storage
M
Min Xie
College of Computer Science and Technology, National University of Defense Technology
Huijun Wu
Huijun Wu
National University of Defense Technology
storage systemsgraph neural networks
Z
Zhenwei Wu
College of Computer Science and Technology, National University of Defense Technology
W
Wenzhe Zhang
College of Computer Science and Technology, National University of Defense Technology