🤖 AI Summary
This work addresses the inefficiency of large language models (LLMs), which rely on a monolithic context window as memory without hierarchical organization, leading to structural redundancy and resource waste. To overcome this limitation, the paper introduces the concept of virtual memory into LLM systems, proposing a novel L1–L3 multi-level memory architecture. A transparent proxy situated between the client and the inference API enables demand paging, page fault detection, and working-set page pinning. Integrated with dialogue compression and a page-fault-driven replacement policy, the system achieves up to a 93% reduction in context memory usage—from 5,038 KB to 339 KB—in real-world production settings, while offline simulations show a page fault rate of only 0.0254%, effectively transcending the constraints of conventional fixed-size context windows.
📝 Abstract
The context window of a large language model is not memory. It is L1 cache: a small, fast, expensive resource that the field treats as the entire memory system. There is no L2, no virtual memory, no paging. Every tool definition, every system prompt, and every stale tool result occupies context for the lifetime of the session. The result is measurable: across 857 production sessions and 4.45 million effective input tokens, 21.8% is structural waste.
We present Pichay, a demand paging system for LLM context windows. Implemented as a transparent proxy between client and inference API, Pichay interposes on the message stream to evict stale content, detect page faults when the model re-requests evicted material, and pin working-set pages identified by fault history. In offline replay across 1.4 million simulated evictions, the fault rate is 0.0254%. In live production deployment over 681turns, the system reduces context consumption by up to 93% (5,038KB to 339KB); under extreme sustained pressure, the system remains operational but exhibits the expected thrashing pathology, with repeated fault-in of evicted content.
The key observation is that the problems the field faces, such as context limits, attention degradation, cost scaling, lost state across sessions, are virtual memory problems wearing different clothes. The solutions exist: working set theory (Denning, 1968), demand paging, fault-driven replacement policies, and memory hierarchies with multiple eviction-managed levels. We describe the architecture of a full memory hierarchy for LLM systems (L1 through persistent storage), report on the first three levels deployed in production use (L1 eviction, L2 fault-driven pinning, L3 model-initiated conversation compaction), and identify cross-session memory as the remaining frontier.