LAST: Looped Audio Spectrogram Transformer

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the parameter redundancy and high computational costs that hinder the deployment of deep Transformers for audio recognition. To overcome these limitations, this work proposes the Recurrent Audio Spectrogram Transformer (R-AST), which introduces a novel low-cost recurrence mechanism. By reusing fixed spectrogram feature blocks and iteratively refining only the class token, R-AST enables efficient deep reasoning. This architecture significantly extends model depth and enhances robustness with minimal overhead. Evaluated on AudioSet, R-AST achieves a mean average precision (mAP) of 0.345, surpassing the 12-layer baseline by 2.1% while reducing the parameter count by 49.4% and increasing throughput by 9.8%.
📝 Abstract
Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.
Problem

Research questions and friction points this paper is trying to address.

Audio Spectrogram Transformer
Computational Efficiency
Model Depth
Audio Recognition
Parameter Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Transformer
Audio Spectrogram Transformer
Class Token Refinement
Parameter Efficiency
Computational Efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haider Al-Tahan
Georgia Institute of Technology
Sean O'Brien
Sean O'Brien
PhD @ UC San Diego (previously UC Berkeley, Meta AI)
natural language processingdecoding methodslarge language modelsdark matter
A
Anastasia Razdaibiedina
Google DeepMind
N
N. Apurva Ratan Murty
Georgia Institute of Technology