Score
Design and specify on-disk file formats whose data is produced only by appending, including byte-level layout, metadata, and durable write semantics. Build compact lookup/index structures and write/compaction procedures, and optimize layouts and access patterns for efficient memory-mapped reads and fast dataset construction and querying.
Existing file systems fail to harness the byte-addressable potential of CXL/PCIe memory-semantic SSDs, resulting in suboptimal read/write performance, severe write amplification, and low storage efficiency. This paper introduces ByteFS—the first file system supporting unified byte- and block-granularity access. Its core innovations are: (1) a hybrid B+-tree index architecture integrating both byte- and block-level abstractions; (2) an on-device DRAM-resident structured logging and data-merge write mechanism implemented in SSD firmware; and (3) a cross-layer consistency caching strategy jointly orchestrating host page cache and SSD-side cache. Experimental evaluation demonstrates that ByteFS achieves up to 2.7× higher throughput and reduces SSD write traffic by 5.1× compared to state-of-the-art NVM/SSD file systems, while strictly guaranteeing crash consistency and persistence semantics.
Log-structured table formats (e.g., Delta Lake, Iceberg, Hudi) in data lakes suffer from excessive small files due to append-only writes and metadata-heavy operations, degrading query performance, increasing storage costs, and limiting system scalability. Existing compaction mechanisms lack flexibility, pursue narrow objectives, and fail to balance benefits against operational overhead. To address this, we propose Scalable Adaptive Compaction (SAC), an extensible, workload-aware, metadata-driven framework featuring dynamic threshold tuning, lightweight online evaluation, and a modular rule engine. SAC is production-deployed via the OpenHouse control plane. Evaluated on LinkedIn’s production workloads and synthetic benchmarks, SAC reduces file counts by up to 92%, improves typical query latency by 3.8×, and maintains bounded runtime overhead.
This study addresses the performance degradation in lakehouse tables caused by small-file accumulation, a problem exacerbated by the absence of principled merge-triggering mechanisms. The authors develop an open simulation framework that generates diverse table layouts based on Apache Iceberg and extracts 17 metadata features to predict post-merge file reduction ratios with high accuracy using XGBoost (R²=0.998). Their work is the first to systematically uncover strong correlations between metadata characteristics and merge efficacy. They further propose a model-free merging strategy relying solely on a single partition-level threshold, which demonstrates remarkable generalization across heterogeneous workloads (R²=0.976). Experimental results show that this approach substantially improves metadata-intensive query performance while also revealing its subtle impact on the parallelism of full-table scan compaction.
Existing LSM-tree and B-tree–based KV stores—as well as systems like FASTER—face bottlenecks in index memory overhead, log-compaction efficiency, and hot working-set management for large-scale KV services characterized by memory constraints, highly skewed access patterns, and datasets far exceeding main memory capacity. Method: This paper proposes a hierarchical record-oriented KV storage architecture featuring (i) a novel multi-threaded lock-free log compaction mechanism, (ii) a lock-free read-modify-write (RMW) algorithm, (iii) a two-level hash index coupled with a read cache, and (iv) tight integration of LSM-style cold-hot data separation with a CPU-optimized lock-free concurrency engine, specifically tailored to NVMe SSD characteristics. Results: Under typical skewed workloads, the system achieves 2.1×–11.75× higher average throughput than RocksDB, significantly reduces write amplification, and delivers low latency with high robustness.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
This work addresses memory exhaustion and I/O bottlenecks in processing petabyte-scale image datasets—such as 1.4 PB electron microscopy volumes or 150 TB organ atlases—by introducing a streaming single-pass architecture based on a sweep execution model. The approach aligns disk reads with a one-dimensional sweep order and combines windowed operations with overlap-aware tiling to enable efficient processing under tight memory constraints. A domain-specific language (DSL) is designed to automatically optimize window sizes, fuse pipeline stages, and schedule multi-pass sweeps at compile time and runtime. The system supports Zarr, HDF5, and slice-based formats without requiring full-image residency in memory, achieving significantly higher throughput, near-linear I/O scaling, and predictable memory usage while seamlessly integrating with existing segmentation and morphological analysis toolchains.
This work addresses the longstanding challenge of reconciling correctness guarantees with high performance in database storage systems by proposing a compositional construction methodology grounded in formal specifications. The approach leverages formal specifications to guide a Java implementation and employs storage equivalence proofs to rigorously ensure correctness. Crucially, it decouples performance optimizations from functional logic, enabling flexible composition of components. Using this methodology, the authors reimplement RocksDB’s tiered storage architecture and develop CobbleDB, a prototype system that demonstrates strong real-world performance and practical applicability while maintaining strict correctness guarantees.
This work presents the first deterministic, fully indexable dictionary (FID) in the standard Word-RAM model that supports worst-case efficient rank, select, and bit update operations. Building upon Pătraşcu–Thorup fusion trees and leveraging a precomputed table access mechanism ℳ_B under the constant-time integer multiplication assumption, the proposed parameterized structure achieves a space usage of lg(ᵘₙ) + O(nw^ε/ε) bits for any ε ≤ 1/2, with rank/select₀ and update operations running in O(1/ε + log_w n) time and select₁ in O(log_w n) time. Notably, setting ε = 1/√lg w reduces the redundancy to O(n log w) bits, thereby achieving—within o(n√w) redundancy—the Fredman–Saks lower bound on operation time simultaneously for all supported queries.
This work addresses the high software overhead and performance bottlenecks of LSM-tree-based databases on fast storage devices, primarily caused by frequent system calls and user-kernel context switches during compaction. To tackle this, the authors propose the first integration of eBPF and io_uring into LSM-tree compaction optimization, migrating critical I/O paths into the kernel to enable zero-intrusion kernel acceleration. This approach significantly reduces software stack overhead without altering the underlying storage format or compaction algorithms. Experimental results demonstrate a 99% reduction in system calls and a 50% decrease in compaction time. Furthermore, under write-intensive workloads, the system achieves a 75% improvement in throughput and a 40% reduction in p99 latency.
This work addresses the limitations of traditional retrieval-augmented generation (RAG) systems, which rely on server-side computation and suffer from privacy risks, high latency, and substantial storage overhead, while on-device deployment is hindered by resource constraints that compromise both retrieval efficiency and generation quality. To overcome these challenges, the authors propose a unified on-device RAG model that jointly optimizes retrieval and context compression for the first time. By sharing document representations across both tasks, the model enables efficient retrieval and compact context generation within a single architecture, eliminating redundancy from multiple separate models. The approach achieves generation performance comparable to conventional RAG using only approximately one-tenth of the context length, while maintaining embedding storage costs no higher than existing multi-vector retrieval methods, thereby significantly enhancing the practicality and efficiency of on-device RAG.