DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "context rot" problem, where large language models suffer sharp performance degradation as input length increases during long-context reasoning. To this end, it proposes a novel Grounding-Reasoning decoupling paradigm that separates evidence retrieval from logical inference. Methodologically, inspired by Spark-like distributed architectures, multiple worker nodes perform parallel context retrieval while a central model handles planning and evidence synthesis. Furthermore, Group Relative Policy Optimization (GRPO) reinforcement learning and atomic task mapping techniques are introduced to optimize training. Experimental results demonstrate that this framework achieves 78.4% accuracy at the million-token scale, surpassing full-context models while reducing inference costs by over 80%.
📝 Abstract
While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exhausts the representational capacity needed for complex reasoning. To resolve this, we propose Grounding-Reasoning Disaggregation via DIStributed long COntext scaling (DISCO). Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding. A central Driver LLM, trained via Reinforcement Learning (GRPO) to optimize planning, orchestrates execution by dynamically mapping queries into atomic extraction tasks and reducing the gathered evidence to synthesize a final answer. By isolating reasoning from raw context noise, DISCO effectively eliminates context rot. On RULER-QA (1M tokens), it maintains 78.4% accuracy where standard baselines collapse. Furthermore, it outperforms full-context models by up to 9.8 points on LongBench v2 and matches frontier models like Gemini-3-Pro-Preview while reducing inference costs by over 80%, establishing a highly efficient paradigm for robust long-context inference.
Problem

Research questions and friction points this paper is trying to address.

context rot
long context scaling
large language models
grounding-reasoning disentanglement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grounding-Reasoning Disaggregation
Distributed Long Context Scaling
Reinforcement Learning (GRPO)
Context Rot
Parallel Localized Grounding