RAMPART: Registry-based Agentic Memory with Priority-Aware Runtime Transformation

📅 2026-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key challenges in large language model (LLM) agents, including inefficient context assembly, unstructured memory management, and insufficient access control during runtime. To overcome these limitations, the authors propose a compile-time-defined, permissioned memory model that introduces a registry of pure memory blocks annotated with ownership tags, enabling block-level operations with zero prompting overhead. Structured context assembly is achieved through five composable primitives—promote, gate, write, evict, and rollback—augmented by provenance labels, non-evictable author flags, and pattern-based eviction policies. Experimental results demonstrate substantial performance gains: block grouping improves task success rates by tens of percentage points; relevance-based gating reduces prompting costs by 67.8% while recovering 83% of success rates; and even small models with the proposed intervention outperform larger, unmodified counterparts.
📝 Abstract
RAMPART is a compile-time memory model and pure in-RAM block registry for LLM-based agents. Context assembly is a programmable runtime operation where content is compiled from a structured registry under explicit policy for ordering, inclusion, and eviction. Five composable primitives (promote, gate, write, evict, rollback) act on named addressable blocks before compilation at zero prompt-token cost. Provenance tags and non-evictable authorship flags implement a permissioned memory model with block-level ownership. Controlled probes with Qwen3-8B Q4 show that compile-time placement and the structural relationship between blocks and the task query affect task success, with the cliff falling at roughly the seventh block position when the task follows the registry and the twelfth when it precedes. Grouping the critical block with content-adjacent neighbours and promoting the group as a unit lifts task success by tens of percentage points at positions where single-block placement fails. Cross-model replication on Qwen2.5-7B, Llama-3.1-8B, Mistral-7B-v0.3, and Qwen3-14B shows the content-priming effect appears at the same absolute positions across families, with magnitude varying with model strength. Block grouping raises Mistral's mean pass rate roughly fivefold at the hardest registry size, and a smaller model with the intervention can outperform a larger model without it in the mid-registry zone. Relevance gating reduces prompt cost by 67.8\% while recovering 83% of the promoted-condition success rate. Schema eviction produces 0% invocations against 100% with the schema present, a property policy-based approaches cannot guarantee by construction. Shared-registry coordination reduces inter-agent communication to a method call at zero coordination token cost.
Problem

Research questions and friction points this paper is trying to address.

memory management
LLM-based agents
structured registry
context assembly
block-level ownership
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic memory
compile-time memory model
block registry
context assembly
zero prompt-token cost
🔎 Similar Papers