Telescopic Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that single language models struggle to accommodate multiple compute budgets without repeated training. To this end, it proposes a continuous elastic training objective coupled with a nested-capacity Transformer architecture. By leveraging random prefix sampling and full-anchor joint supervision, the selection of inference operating points is shifted from architectural design to training-time configuration, enabling valid outputs at arbitrary depths without additional inference overhead. Experimental results demonstrate that the proposed method reduces the area under the quality-budget curve by 43–44% and decreases GPU costs by approximately 12%, while maintaining full-capacity performance on par with standard baselines.
📝 Abstract
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.
Problem

Research questions and friction points this paper is trying to address.

Language Models
Compute Budgets
Nested Transformer
Elastic Models
Fixed-exit Suites
Innovation

Methods, ideas, or system contributions that make the work stand out.

Telescopic Language Model
stochastic prefix supervision
nested-capacity Transformer
compute budget continuum
elastic language model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhilin Guo
University of Cambridge
B
Boqiao Zhang
University of Cambridge
Hakan Aktas
Hakan Aktas
University of Cambridge
RoboticsDeep LearningAI
Kyle Fogarty
Kyle Fogarty
University of Cambridge
Geometry ProcessingGeometric Deep LearningApplied Mathematics
N
Nursena Koprucu Aslan
University of Cambridge
W
Wenzhao Li
University of Cambridge
C
Canberk Baykal
University of Cambridge
A
Albert Miao
University of Cambridge
S
Siyu Hong
University of Cambridge
Y
Yixiao Liu
University of British Columbia
A
Adam Wu
University of Cambridge
Ashish Kumar Singh
Ashish Kumar Singh
Google
S
Sakar Khattar
Google
Chenliang Zhou
Chenliang Zhou
University of Cambridge
machine learninggenerative artificial intelligencecomputer visioncomputer graphics
W
Weihao Xia
University of Cambridge
C
Cristina Nader Vasconcelos
Google
C
Cengiz Oztireli
University of Cambridge, Google