🤖 AI Summary
This study addresses the limitation of existing eviction strategies in fixed-capacity streaming video memory, which overlook the cumulative impact of repeated updates on future predictive states. To this end, this work introduces the concept of "predictive states" and a counterfactual eviction framework that formulates memory eviction as counterfactual control within the predictive space. By leveraging a frozen multi-horizon JEPA model to evaluate future representational utility under different actions, the approach dynamically balances error correction and adaptation to optimize retention policies, thereby suppressing predictive drift without online parameter updates. Furthermore, this paper presents the MABS-Bench benchmark. Extensive experiments across multiple video datasets demonstrate that the proposed method significantly reduces predictive drift and enhances task performance, consistently outperforming existing baselines.
📝 Abstract
Fixed-capacity streaming video memory requires repeated eviction decisions whose effects accumulate over time. Yet existing policies are evaluated primarily in terms of retained information or downstream accuracy, leaving how repeated updates alter the futures supported by memory largely unexamined. We define a memory's predictive state as the future representations supported by its retained history and formulate eviction as counterfactual control over transitions in this space. We introduce SPACE (Sparse Predictive Attractor via Counterfactual Eviction), which uses a frozen multi-horizon JEPA to predict the future representations induced by alternative eviction actions. Counterfactual utility identifies future-useful alternatives, while slow predictive-basin geometry determines when to correct avoidable drift and when to adapt to sustained predictive change, without online parameter updates. We further introduce MABS-Bench, which evaluates future-task sufficiency, within-regime predictive stability, transition responsiveness, and perturbation recovery under matched causal streams and memory budgets. Across multiple video datasets, SPACE yields consistent improvements in dataset-native task performance while reducing predictive-state drift.