Sparse Token Routing in Efficient Transformers

📅 2026-08-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究使用SEWN模型测试了Transformer中令牌稀疏路由的有效性,通过学习门控机制将令牌分配到轻量级或全容量处理,验证了不同令牌需要不同计算资源的假设。
📝 Abstract
Efficient-transformer research often motivates token pruning and adaptive computation with the claim that not all tokens require equal computational effort. We test this claim end to end using SEWN, a two-stream Transformer that routes tokens through either lightweight or full-capacity processing using a learned gate. Across our experiments, routing introduces negligible accuracy change relative to parameter-matched baselines, while the gate's token-importance signal depends critically on how it is learned. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without changing task accuracy.
Problem

Research questions and friction points this paper is trying to address.

Sparse Token Routing
Efficient Transformers
Token Pruning
Adaptive Computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Token Routing
Efficient Transformers
Two-Stream Transformer
Learned Gate
Token-Importance Signal
💼 Related Jobs
No related jobs found.
S
Sai Krishna Arthanari
Institute for Artificial Intelligence and Data Science (IAD), University at Buffalo
J
JaeHyeong Chang
Institute for Artificial Intelligence and Data Science (IAD), University at Buffalo
C
Chengzhe Sun
Institute for Artificial Intelligence and Data Science (IAD), University at Buffalo
S
Siwei Lyu
Institute for Artificial Intelligence and Data Science (IAD), University at Buffalo