Score
Designs, implements, and evaluates systems and policies that govern the recording, encoding, indexing, retrieval, transfer, updating, and pruning of memory or data across multiple tiers (e.g., short-term buffers, intermediate caches, long-term repositories). Includes lifecycle modeling and operational procedures to decide which experiences or records to persist, how to maintain and compact growing stores, and how to manage buffers and movement between hierarchy levels to support reuse and rapid adaptation.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
This work addresses the joint optimization of media selection, capacity allocation, and data placement (replication vs. tiering) for key-value caching across heterogeneous NVM/DRAM/disk storage under memory budget constraints. We introduce the first systematic modeling framework for multi-level non-volatile cache configurations, analytically characterize the operational regimes where replication or tiering dominates, and propose an adaptive configuration policy grounded in device failure rates and data update frequencies. Our methodology integrates cache access behavior modeling, hierarchical configuration optimization, and empirical validation using memcached benchmarks. Results demonstrate that tiering substantially outperforms replication under low device failure rates and high update workloads. Key contributions include: (1) a deployable, low-overhead configuration algorithm; (2) quantitative design guidelines for heterogeneous cache deployment; and (3) theoretical foundations for the reliability–performance trade-off in tiered caching systems.
Existing memory tiering systems rely on manually tuned, fixed thresholds, making it difficult to sustain high performance across diverse workloads. This paper proposes a parameter-free adaptive memory tiering mechanism. First, it designs short- and long-term moving-average hot-page detectors to dynamically classify page temperature. Second, it formulates a cost-benefit–driven migration decision model that eliminates threshold dependency. Third, it introduces a bandwidth-aware batch scheduling policy to improve I/O efficiency. Deeply integrated into the kernel’s memory management subsystem, the solution requires no configuration and delivers out-of-the-box usability. Evaluated across multiple workload classes, it achieves over 97% of the performance of the best manually tuned baseline, while outperforming the unoptimized baseline by 1.26×–2.3×. The approach significantly enhances robustness and practicality in real-world deployment.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
This work addresses a critical limitation in existing memory tiering systems, which often neglect the significant impact of page size and hardware topology on performance when making page migration decisions. To overcome this, the authors propose a lightweight, eBPF-based page admission control mechanism that dynamically integrates page size and hardware topology information through a working-set-agnostic page profiling approach, enabling user-customizable migration policies. Notably, this solution requires no kernel modifications and seamlessly integrates into current systems. Evaluation across three mainstream memory tiering frameworks and seventeen diverse workloads demonstrates substantial performance gains, with geometric mean throughput improvements of up to 17.7% and peak speedups reaching 75% in specific scenarios.
This work addresses the significant yet underexplored impact of parameter configuration on performance in memory tiering systems, where efficient automated tuning mechanisms are lacking. The authors propose PTMT, a lightweight framework that systematically categorizes memory tiering parameters and reveals their performance sensitivity for the first time. PTMT introduces a hybrid adaptive tuning mechanism that combines offline performance profiling with online reinforcement learning to achieve high-efficacy optimization at low overhead. By co-designing memory access profiling and page migration, PTMT improves performance by 30%, 26%, 21%, and 14% over TPP, UPM, Colloid, and AutoNUMA, respectively, and outperforms the best existing approaches by 32% on average.
This work addresses the challenge of knowledge updating in large language models, which typically necessitates costly retraining. Building upon the Compositional Multi-layer Memory (CMM) architecture, the authors propose a version-aware operational layer that compiles high-level semantic edits into ordered, composable memory primitive transactions. By introducing versioned CMM and transactional CMM, knowledge modifications are modeled as reversible and reusable structured transactions, enabling fine-grained replacement, rollback, historical tracing, and localized updates. This approach substantially reduces reliance on full model retraining while ensuring editing efficiency and traceability.
This study addresses the challenge of coordinating multiple autonomous mobile robots (AMRs) in high-density manufacturing environments to jointly execute storage, retrieval, and reshuffling tasks with time windows. It is the first to incorporate dynamically arriving storage tasks into a buffer job scheduling framework. The authors propose a scalable hierarchical heuristic: an upper layer optimizes the sequence of unit-load tasks using A* search, while a lower layer employs constraint programming for real-time multi-robot coordination. A binary integer programming model is also developed to serve as an exact solution benchmark. Experimental results demonstrate that the proposed approach achieves solution quality comparable to the exact method while reducing computation time by several orders of magnitude, thereby substantially enhancing the feasibility of real-time control in high-density scenarios.
This work challenges the prevailing assumption that large language model (LLM) agents inherently function without structured long-term memory, presenting the first systematic investigation into file systems as a memory substrate. The authors introduce a unified memory architecture comprising three collaborating agent types—manager, searcher, and executor—that jointly maintain a directory-tree-based Markdown storage system augmented with sandboxed shell tools, dedicated memory interfaces, and chunked retrieval strategies. Through comprehensive evaluation, they assess how organizational schemes, tooling, and model capabilities jointly influence memory coherence and task performance. Results demonstrate that well-structured memory organization can reduce retrieval costs by approximately 50%, yet most agents struggle to sustain such structure over time. Notably, changes to the toolset exert an impact on memory morphology comparable to switching the underlying LLM, underscoring that memory architecture constitutes a design space rather than a fixed default.
Traditional caching policies such as LRU and LFU fail in semantic retrieval scenarios due to the absence of temporal locality and skewed access frequency. This work formulates semantic cache management as an online replacement problem with switching costs and introduces SOLAR, a novel framework that determines update时机 through cumulative regret and leverages Bayesian online learning to optimize content selection from implicit feedback. Theoretical analysis establishes that SOLAR achieves a constant competitive ratio (≤3) and a near-optimal regret bound of O(√(KT log T)), thereby overcoming the dependence on cache size and time horizon inherent in conventional approaches. Empirical results demonstrate that SOLAR improves performance by 5%–75% over FIFO under tight cache constraints, reveals for the first time that classical caching strategies systematically underperform FIFO under semantic workloads, and uncovers an inverted U-shaped relationship between retrieval pool size and quality.