KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决语言模型扩展中的计算成本问题,提出KV-Invariant Transformer Expansion(KITE)方法,通过参数重分配减少训练与推理成本。
📝 Abstract
Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling to a larger model, while ensure the larger model indeed outperforms smaller baselines. We introduce KV-Invariant Transformer Expansion (KITE), a scaling paradigm that achieves this goal. It trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV. Consequently, during inference, prefilling KV only relies on the smaller part of the model, so the inference costs are saved. As a concrete instantiation, we present Step Scale Transformer (SST), a two-tower decoder in which one tower produces KV and the other reads them. At comparable cumulative training compute, SST, a 67B MoE model with 2.15B active body parameters per decode token, achieves lower training loss than 47B and 63B MoE Transformers with 1.48B and 2.02B active body parameters, respectively, while reducing estimated inference cost by 6.7% and 31.6%.
Problem

Research questions and friction points this paper is trying to address.

Scaling
Computation Costs
Model Quality
KV-Invariant
Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-Invariant
Transformer Expansion
Efficient Scaling
Step Scale Transformer
Inference Cost Reduction
🔎 Similar Papers
No similar papers found.