🤖 AI Summary
This study addresses the challenge of retaining and recovering cross-session multimodal evidence under constrained budgets in long-horizon tasks by proposing a compact and traceable multimodal memory organization framework. Methodologically, it constructs a bounded active index that integrates redundancies while preserving complementary records through relation-aware updates, coupled with a budgeted routing mechanism to dynamically expand source evidence at query time. The proposed approach effectively distinguishes temporally and semantically incompatible observations, achieving reliable retention of fine-grained visual cues and provenance links under a fixed reader budget. Ultimately, this work provides critical support for downstream reasoning by ensuring both memory compactness and evidential traceability across extended task horizons.
📝 Abstract
Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records. At query time, budgeted routing selects useful index pages and expands their associated source evidence under a fixed reader budget. Together, these mechanisms establish a compact, provenance-preserving multimodal memory organization for cross-session long-horizon tasks, retaining temporal distinctions and source links required for reliable downstream reasoning. Code is available at https://github.com/HuzhouNLP/C3M.