🤖 AI Summary
This study addresses the challenge recurrent models face in balancing persistent memory retention against computational overhead during delayed state retrieval. To this end, we propose a recurrent memory architecture that integrates chunk-local attention with pattern indexing. Specifically, the method introduces learnable pattern embeddings as shared references, enabling efficient read and write operations through a feedback-free writing mechanism coupled with a boundary-commit aggregation strategy. Experimental results demonstrate that the proposed architecture significantly improves write-value retention and state-update recovery under long-delay scenarios. Ultimately, this work establishes a new optimal trade-off between memory persistence and computational efficiency for recurrent architectures.
📝 Abstract
Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-based recurrent memory architecture combining chunk-local attention with a persistent, schema-indexed phase state. Learned schema embeddings provide a shared representational reference for reading and writing. Reads use the current state, whereas writes use the layer input and static schema embeddings, excluding direct feedback from that layer's own state. Chunk-boundary commits aggregate bounded phase increments through forward computation. The same parameters also support full-history attention training before and during recurrent training. We studied selective updates, preservation, and delayed retrieval in a controlled address--value task, comparing three-layer models with approximately matched parameter counts and persistent-state dimensions. Across nine address/value settings and three training seeds, SchemaMem has higher mean written-value retention at four times the maximum training delay than both baselines, which are trained toward a higher in-range accuracy target. Updated-value recovery favors SchemaMem in all nine settings against Mamba-3 and seven against Gated DeltaNet. Defaults consistently favor Gated DeltaNet over SchemaMem at that delay, and SchemaMem requires substantially more optimization steps. These results identify a promising retention--optimization trade-off in schema-indexed recurrence.