spectral token compression

Designs and evaluates algorithms and systems that compress tokenized representations by transforming tokens into spectral (frequency-domain) components, identifying and removing redundant spectral content, and producing reduced token sequences via pooling, reduction, or discrete semantic encoding. Builds methods that preserve semantic information and downstream-task accuracy while minimizing token count or bit-rate, and analyzes trade-offs between spectral filtering, token-centric compression, and discrete semantic communication.

spectraltokencompression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Discrete Audio Tokens: More Than a Survey!

Jun 12, 2025
PM
Pooneh Mousavi
🏛️ Concordia University | Mila-Quebec AI Institute | The Hebrew University of Jerusalem | Carnegie Mellon University | Microsoft | Université de Toulon | Google | Apple | Laval University | University of Cambridge | University of Illinois at Urbana-Champaign | National Taiwan University | Université de Montréal

Existing discrete audio tokenization research lacks unified, cross-task and cross-domain evaluation. Method: We systematically survey and benchmark state-of-the-art methods across speech, music, and general audio domains, proposing the first comprehensive taxonomy spanning codec architecture, quantization mechanisms, training paradigms, streaming support, and application dimensions. We design a multi-objective joint optimization framework integrating reconstruction loss, semantic fidelity, and LLM alignment, unifying VQ/RVQ, GAN/MAE, and streaming token generation techniques. Contribution/Results: We conduct horizontal evaluation and controlled ablation studies across 12 standardized benchmarks, identifying critical bottlenecks. We open-source a standardized tokenizer database and core results, establishing—for the first time—the empirical trade-off boundary among reconstruction quality, inference latency, and generalization capability.

Analyze trade-offs and highlight open challengesEvaluate tokenizers on reconstruction and downstream tasksSystematic review and benchmark of discrete audio tokenizers

A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication

Sep 18, 2025
RC
Ryan Collette
🏛️ Systems & Technology Research

Existing speech codecs do not explicitly disentangle semantic hierarchies, making it challenging to simultaneously preserve perceptual quality and downstream task performance at ultra-low bitrates (e.g., <1.5 kbps). To address this, we propose the first decoupled framework for semantic speech compression, introducing hierarchical semantic representations—explicitly separating and differentially encoding phonetic, prosodic, emotional, and speaker-related features—derived from generative speech models into the codec architecture. Leveraging a semantic communication paradigm and multi-granularity reconstruction, our method achieves or surpasses state-of-the-art performance of codecs such as EnCodec on automatic speech recognition, emotion analysis, and speaker verification, while operating at 2–4× lower bitrates. Crucially, intelligibility and naturalness are preserved. This work establishes a novel paradigm for ultra-low-bitrate semantic speech communication.

Achieving ultra-low bandwidth voice communication without quality lossLeveraging semantic representations to reduce bitrates significantlyOutperforming existing codecs in perceptual quality and task performance

dMel: Speech Tokenization made Simple

Jul 22, 2024
RH
Richard He Bai
🏛️ Apple

Existing speech tokenization methods rely on audio compressors, incurring high computational overhead and poor cross-domain generalization. This paper proposes dMel—a training-free, streaming-capable, and robust discrete speech representation—achieved by intensity-based binning of energy per frequency band in log-Mel spectrograms, enabling lightweight tokenization. Its core innovation lies in the first unified optimization of text-to-speech (TTS) and automatic speech recognition (ASR) within a single LM-style Transformer architecture; it employs parallel high-dimensional token encoding/decoding to jointly balance efficiency and representational capacity. Experiments demonstrate that dMel matches or surpasses task-specific models in both synthesis and recognition performance, while significantly reducing computational complexity and deployment barriers. By eliminating the need for auxiliary neural compressors and dedicated training, dMel establishes a simple, efficient, and scalable representation paradigm for speech foundation models.

Enabling unified modeling for speech synthesis and recognitionImproving robustness to out-of-domain audio signalsSimplifying speech tokenization for effective language modeling

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Aug 29, 2024
SJ
Shengpeng Ji
🏛️ Zhejiang University | Alibaba Group | Meta

To address the low compression efficiency and poor fidelity of discrete quantization for high-dimensional audio signals, this paper introduces WavTokenizer—the first efficient discrete encoder-decoder specifically designed for audio language modeling. It innovatively constructs a wide vector-quantized (VQ) codebook space, integrated with an extended-context Transformer, multi-scale GAN discriminators, and an inverse short-time Fourier transform (iSTFT)-based reconstruction architecture, enabling compression of 1-second 24 kHz audio into only 40–75 semantically rich tokens. WavTokenizer achieves state-of-the-art performance across speech, music, and general audio reconstruction: it attains a new industry-leading UTMOS score (+0.32), improves VQ codebook utilization by 37%, and significantly enhances both perceptual quality and semantic consistency. The model is lightweight, open-source, and natively compatible with downstream audio generation tasks.

Efficient audio signal compressionEnhancing subjective audio qualityOptimizing acoustic codec tokenizer

Latest Papers

What's happening recently
View more

Existing acceleration methods for diffusion models often compromise classification capability when compressing computation, struggling to balance generation quality and discriminative performance. This work proposes BiGain—a training-free, plug-and-play framework that jointly enhances both generative and classification performance in accelerated diffusion for the first time. At its core lies a spectrum-aware dual-operator compression mechanism: Laplacian gating enables token merging, while interpolation-extrapolation-based KV downsampling preserves high-frequency details and maintains low- and mid-frequency semantics. On ImageNet-1K, BiGain achieves a 7.15% absolute improvement in classification accuracy and a 0.34 reduction in FID (a 1.85% relative improvement) under a 70% token compression rate, significantly advancing the trade-off between speed and accuracy.

accelerationclassification accuracydiffusion models

This work addresses the unclear trade-offs among compression efficiency, structural inductive bias, and cross-domain robustness in large language model tokenizers. Viewing tokenization through an information-theoretic lens as structured compression, the authors propose a variant of Byte Pair Encoding (BPE) integrated with principles from compressed sensing and introduce metrics such as channel capacity utilization. They systematically analyze how vocabulary size and training data volume influence text entropy distribution and contextual predictability. Experimental results reveal that while increasing training data enhances token diversity, it simultaneously strengthens contextual predictability. The proposed framework effectively quantifies tokenizer performance, offering both theoretical grounding and practical guidance for designing general-purpose, compression-oriented tokenization strategies and downstream modeling.

compression efficiencydomain shiftlarge language models

This study addresses the computational inefficiency in time series language models arising from the unified treatment of time series and prompt tokens, which exhibit fundamentally different information structures. The work reveals an asymmetry in token importance: time series tokens contribute unevenly across the frequency spectrum, while the influence of prompt tokens diminishes with model depth. Leveraging this insight, the authors propose an adaptive token compression framework that dynamically compresses time series tokens via spectral analysis and progressively prunes prompt tokens in deeper layers, enabling hierarchical, frequency-aware, non-uniform budget allocation. Evaluated across forecasting, classification, imputation, and anomaly detection tasks, the method achieves up to 7.68× speedup and improves performance in 78% of experimental settings.

asymmetric tokenstime series language modelstoken compression

Existing neural audio codecs struggle to simultaneously preserve acoustic fidelity and capture semantic content, and incorporating semantic information often degrades reconstruction quality. To address this challenge, this work proposes STACodec, a unified codec framework that introduces a Semantic Token Allocation (STA) mechanism to inject semantic representations from self-supervised models into the first layer of Residual Vector Quantization (RVQ). Furthermore, a Semantic Pre-Distillation (SPD) module is designed to directly predict semantic tokens during inference without relying on external semantic tokenizers. This approach achieves a synergistic optimization of acoustic detail and semantic capability, maintaining high reconstruction fidelity while significantly enhancing performance on downstream semantic tasks.

acoustic fidelityaudio codecsneural audio compression

Hot Scholars

JL

Jinghui Lu

ByteDance Inc., School of Computer Science, University College Dublin
Natural Language ProcessingMulti-ModalityLLMHuman-in-the-loop Learning
JK

Jan Kautz

Vice President of Research, NVIDIA Research
Computer VisionMachine LearningVisual Computing
PM

Pavlo Molchanov

NVIDIA Research
AIMachine LearningEfficient Deep LearningSemi-supervised learning
PH

Pengcheng Huang

Computer Engineering Group, ETH Zurich
Intelligent Learning SystemsCyber Physical Systems
ZD

Zhicheng Dou

Renmin University of China
Information RetrievalRetrieval Augmented GenerationLarge Language ModelsGenerative IR