🤖 AI Summary
This work addresses the lack of systematic evaluation tools for memory write strategies under byte-budget constraints and data stream drift. We propose the first comprehensive benchmarking framework specifically designed for byte-constrained memory write policies. The framework features a controllable non-stationary task generator that simulates document or API drift, an explicit external memory operation interface, a byte-accurate cost model, and standardized performance metrics. It enables reproducible and comparable unified evaluation of diverse write strategies in terms of task success rate and memory efficiency under dynamic data streams, thereby filling a critical gap in benchmarking infrastructure for this emerging research area.
📝 Abstract
We introduce WritePolicyBench, a benchmark for evaluating memory write policies: decision rules that choose what to store, merge, and evict under a strict byte budget while processing a stream with document/API drift. The benchmark provides (i) task generators with controlled non-stationarity, (ii) an explicit action interface for external memory, (iii) a byte-accurate cost model, and (iv) standardized metrics that measure both task success and budget efficiency.