Score
Designs algorithms, rules, and system components that decide which items to remove from bounded storage or resource pools under capacity pressure, specifying eviction criteria, thresholds, and policies as well as partitioning or virtualization of the pool and techniques to preserve locality. Builds simulations and tests to evaluate behavior under pressure and concurrency and to validate correctness and performance of eviction under concurrent access.
This work addresses the challenge in hierarchical caching networks where conventional request-time-based eviction policies fail to assess the impact of removing aligned storage blocks on the connectivity of downstream critical services, often causing service disruptions. The authors model aligned eviction as a weighted vertex separation problem on a graph, precisely computing the downstream demand cut of candidate blocks to reject evictions that compromise protected paths and selecting the feasible eviction with minimal impact. They introduce a novel verifiable service-cut certificate mechanism that, for the first time, unifies capacity reclamation, path continuity, and distributed failure within a certifiable interface, and prove that strategies relying solely on historical information can incur unbounded single-step damage. Experiments across 144 scenarios processing 582.9 trillion packets (404.86 PiB) validate theoretical predictions, reveal a zero-impact extremal phase transition point, and enable full impact vectors and audit samples within supplementary material budgets.
Traditional Linux page caching employs fixed replacement policies, which struggle to adapt to diverse application workloads; kernel-level policy customization remains impractical due to high development complexity and security risks. This paper introduces CacheBPF—the first eBPF-based programmable page cache framework—that safely injects user-defined eviction policies into the kernel via tracepoints, requiring no kernel source modification and supporting cross-process policy sharing with strict isolation. It pioneers the integration of eBPF into core OS memory management, leveraging a policy sandbox and application-aware metadata tagging to ensure both safety and flexibility. Evaluated under realistic workloads, CacheBPF achieves up to 70% higher throughput and 58% lower tail latency compared to baseline policies, demonstrating that fine-grained, workload-specific eviction strategies significantly improve performance for heterogeneous applications.
This work addresses the limitations of traditional Linux page cache eviction policies, which rely on fixed heuristics and struggle to adapt to diverse workloads, thereby constraining cache efficiency. The authors propose the first integration of a lightweight single-layer perceptron directly into the kernel’s page cache subsystem, leveraging eBPF to enable low-overhead, real-time intelligent eviction decisions. The model is trained on kernel-level data collected from real-world workloads to predict page reuse times and dynamically select eviction candidates. Experimental results across a range of representative workloads demonstrate that, compared to a FIFO baseline, the approach improves cache hit rates by up to 10% and achieves a median AUC of 80%, confirming the feasibility and superiority of machine learning–driven cache management within the kernel.
This study addresses the security challenges institutions face in governing the admission, oversight, and revocation of high-privilege external software and AI artifacts—such as dependency packages, container images, and model artifacts. The authors propose the “Custodial Envelope Threshold” framework, which for the first time links privilege scope to the custodial closure of artifacts. It enforces tiered admission control based on execution privileges, permitting infrastructure access only when artifacts satisfy closure conditions regarding identity, provenance path, and revocability. Integrating reference monitors, the principle of least privilege, and transaction cost economics, the framework yields an actionable four-condition sequential assessment tool. Empirical validation via a reference monitor model and a deterministic prediction function demonstrates that high-scrutiny organizations exhibit a pronounced preference for stronger custodial closure when managing high-privilege artifacts across diverse contexts, including package dependencies, GitHub Actions, container images, Terraform modules, and open-source models.
This work addresses the challenge of performance isolation in multi-tenant storage systems, where traditional fair-sharing mechanisms suffer from high preemption latency, leading to severe tail latency spikes. To overcome this, the authors propose the Delta Fair Sharing family of algorithms, which introduces δ-fairness and δ-Pareto efficiency to enable bounded-latency fair sharing on latency-sensitive resources such as write buffers and read caches. The approach strictly confines tail latency spikes for well-behaved clients within δ time units while preserving high resource utilization. Implemented in FAIRDB—a RocksDB-based system—the algorithm supports resource scheduling with explicit latency bounds. Experimental results demonstrate that FAIRDB significantly outperforms existing solutions in isolating interference from high-load tenants, effectively safeguarding the performance of normal clients.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
This work addresses the challenge of GPU memory inflation, request eviction, and throughput degradation in large language model serving under high concurrency, where growing key-value caches strain system resources. The authors develop a discrete-time dynamical model to characterize request admission, memory growth, and eviction mechanisms under continuous batching, revealing for the first time that service-induced congestion is a structurally unstable phenomenon. Theoretical analysis demonstrates that under homogeneous workloads, the no-eviction equilibrium is unstable, often driving the system toward a limit cycle with up to 50% throughput loss. In contrast, heterogeneous workloads with mutually coprime decoding lengths can achieve stable memory-constrained operation through desynchronization. Building on this insight, the study establishes stability criteria and scheduling design principles for multi-class requests, integrating methods from discrete dynamical systems, survival polynomials, and number theory.
This work addresses the lack of systematic evaluation tools for memory write strategies under byte-budget constraints and data stream drift. We propose the first comprehensive benchmarking framework specifically designed for byte-constrained memory write policies. The framework features a controllable non-stationary task generator that simulates document or API drift, an explicit external memory operation interface, a byte-accurate cost model, and standardized performance metrics. It enables reproducible and comparable unified evaluation of diverse write strategies in terms of task success rate and memory efficiency under dynamic data streams, thereby filling a critical gap in benchmarking infrastructure for this emerging research area.