What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation in large language model inference where per-token computational costs are treated uniformly, disregarding inherent difficulty variations and lacking metrics for the actual computation required by individual tokens. We introduce the first formal definition of an upper bound on sufficient token computation, revealing that a small fraction of tokens consumes the majority of compute resources. Methodologically, we construct a hybrid agent system comprising multi-family, cross-capacity models, integrating speculative decoding with dynamic routing to precisely identify high-cost tokens and optimize resource allocation. Experimental results demonstrate that our approach significantly reduces latency by 33% and draft token usage by 32.6% while preserving generation accuracy. This work provides both theoretical grounding and a practical technical pathway for efficient LLM inference.
📝 Abstract
Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token's sufficient compute, which upper-bounds what the token requires. On three core benchmarks, a 0.5B agent reproduces 92--95\% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10\% account for 64--80\% of estimated FLOPs. On all 500 MATH-500 problems, the MoA-derived map helps model routing reduce projected latency from 7.59 to 5.12 seconds while slightly improving accuracy, relative to the best confidence-routing baseline. The MoA-map helps drafting use 32.6\% fewer draft tokens and approximately 20\% lower projected latency than fixed-window drafting at similar accuracy. These comparisons reveal remaining allocation headroom, motivating controllers that exploit sufficient-compute structure.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Agents
sufficient per-token compute
model routing
speculative decoding
compute allocation