Periscope: Extending Frozen Language Models Beyond Their Context Window

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of context windows and accuracy degradation in long-text reasoning for language models by proposing a training-free, grid-based chunked inference method. The approach maps text chunks onto a two-dimensional grid, leveraging local and strided probes to construct evidence graphs rather than processing full contexts. Combined with log-odds scoring, this strategy achieves sub-linear complexity scaling at the square-root level. Without requiring fine-tuning or modifying frozen large language models, the method overcomes hardware memory bottlenecks, enabling single-GPU processing of up to 4.5 million tokens. Experimental results demonstrate superior performance on benchmarks such as LongBench v2, substantially reducing memory requirements while effectively enhancing long-context reasoning capabilities.
📝 Abstract
A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.
Problem

Research questions and friction points this paper is trying to address.

long-context understanding
context window extension
frozen language models
quadratic complexity
evidence retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free Inference
Context Window Extension
Evidence Map
Frozen Language Models
Memory-efficient Attention
🔎 Similar Papers
M
Mohamed Eltahir
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
A
Anas Obayd
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
R
Raed Rashid
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
A
Abdulrahman Alghamdi
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
A
Abdulrahman Mousa
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
A
Abdallah Ahmed
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
Tanveer Hussain
Tanveer Hussain
Lecturer at Department of Computer Science, Edge Hill University
Computer VisionVideo SummarisationSaliency DetectionFire/Smoke Detection
N
Naeemullah Khan
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia