GTR: Gated Token Recurrence for Efficient Dense Prediction

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决自注意力机制在高分辨率图像上的效率问题,本文提出Gated Token Recurrence (GTR)方法,采用无softmax的循环结构,提高了密集预测任务的效率。
📝 Abstract
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. GTR is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment through a linear projection and squared $\ell_2$ loss, without masked-token prediction or intermediate-layer supervision. With Objects365 detector pre-training, GTR-L achieves 58.9 box AP on COCO \texttt{val2017} with 1.908\,ms median batch-one latency under compiled FP16 execution on an RTX~4090. The same backbone also transfers to instance segmentation, pose estimation, oriented detection, semantic segmentation, and monocular depth estimation. In an isolated kernel benchmark, our specialized chunkwise CUDA operator is $4.0\times$ faster than FLA v0.5.0 at 1.6K tokens on RTX~4090. TensorRT deployment on DRIVE AGX Thor achieves 2.282--8.769\,ms median batch-one latency across the evaluated models. These results show that recurrent token mixing can provide an efficient alternative to global softmax attention for high-resolution dense prediction and edge deployment.Project page: https://intellindust-ai-lab.github.io/projects/GTR/
Problem

Research questions and friction points this paper is trying to address.

self-attention
dense prediction
computational cost
image resolution
efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gated Token Recurrence
gated linear attention
spatially enhanced SwiGLU blocks
efficient dense prediction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhe Feng
Didi International Business Group; Intellindust AI Lab; Institute of Automation, Chinese Academy of Sciences
L
Longfei Liu
Intellindust AI Lab
W
Wei Liu
Didi International Business Group
K
Kai Chen
Didi Research
J
Jiangjiang Kong
Didi International Business Group
W
Wei Zhou
Didi International Business Group
Y
Yifeng Qian
Didi Research
Dexiong Chen
Dexiong Chen
Max Planck Institute of Biochemistry
machine learningbioinformaticsgraph machine learningkernel methods
Xuanlong Yu
Xuanlong Yu
Paris-Saclay University & ENSTA Paris, France
Computer VisionDeep LearningUncertainty Estimation
Xi Shen
Xi Shen
Chief Scientist, Intellindust
Deep LearningComputer Vision