Score
Designs and evaluates algorithms and systems that compress tokenized representations by transforming tokens into spectral (frequency-domain) components, identifying and removing redundant spectral content, and producing reduced token sequences via pooling, reduction, or discrete semantic encoding. Builds methods that preserve semantic information and downstream-task accuracy while minimizing token count or bit-rate, and analyzes trade-offs between spectral filtering, token-centric compression, and discrete semantic communication.
Multimodal large language models (MLLMs) face significant efficiency bottlenecks when processing long-context inputs—such as high-resolution images, lengthy videos, and extended audio—due to the quadratic computational complexity of self-attention. To address this, we propose the first unified taxonomy for multimodal long-context token compression, jointly organized along two dimensions: (i) modality-specific redundancy characteristics (spatial, spatio-temporal, and spectral), and (ii) technical mechanisms (transform-based, similarity-based, attention-based, and query-based paradigms). We systematically survey existing approaches, identify core challenges—including fidelity preservation, cross-modal alignment, and adaptive compression—and outline promising future research directions. Furthermore, we publicly release a dynamically updated multimodal compression knowledge base. This work establishes the first structured theoretical foundation and practical guideline for efficient multimodal long-context modeling.
Existing discrete audio tokenization research lacks unified, cross-task and cross-domain evaluation. Method: We systematically survey and benchmark state-of-the-art methods across speech, music, and general audio domains, proposing the first comprehensive taxonomy spanning codec architecture, quantization mechanisms, training paradigms, streaming support, and application dimensions. We design a multi-objective joint optimization framework integrating reconstruction loss, semantic fidelity, and LLM alignment, unifying VQ/RVQ, GAN/MAE, and streaming token generation techniques. Contribution/Results: We conduct horizontal evaluation and controlled ablation studies across 12 standardized benchmarks, identifying critical bottlenecks. We open-source a standardized tokenizer database and core results, establishing—for the first time—the empirical trade-off boundary among reconstruction quality, inference latency, and generalization capability.
Existing speech codecs do not explicitly disentangle semantic hierarchies, making it challenging to simultaneously preserve perceptual quality and downstream task performance at ultra-low bitrates (e.g., <1.5 kbps). To address this, we propose the first decoupled framework for semantic speech compression, introducing hierarchical semantic representations—explicitly separating and differentially encoding phonetic, prosodic, emotional, and speaker-related features—derived from generative speech models into the codec architecture. Leveraging a semantic communication paradigm and multi-granularity reconstruction, our method achieves or surpasses state-of-the-art performance of codecs such as EnCodec on automatic speech recognition, emotion analysis, and speaker verification, while operating at 2–4× lower bitrates. Crucially, intelligibility and naturalness are preserved. This work establishes a novel paradigm for ultra-low-bitrate semantic speech communication.
Existing speech tokenization methods rely on audio compressors, incurring high computational overhead and poor cross-domain generalization. This paper proposes dMel—a training-free, streaming-capable, and robust discrete speech representation—achieved by intensity-based binning of energy per frequency band in log-Mel spectrograms, enabling lightweight tokenization. Its core innovation lies in the first unified optimization of text-to-speech (TTS) and automatic speech recognition (ASR) within a single LM-style Transformer architecture; it employs parallel high-dimensional token encoding/decoding to jointly balance efficiency and representational capacity. Experiments demonstrate that dMel matches or surpasses task-specific models in both synthesis and recognition performance, while significantly reducing computational complexity and deployment barriers. By eliminating the need for auxiliary neural compressors and dedicated training, dMel establishes a simple, efficient, and scalable representation paradigm for speech foundation models.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
To address the low compression efficiency and poor fidelity of discrete quantization for high-dimensional audio signals, this paper introduces WavTokenizer—the first efficient discrete encoder-decoder specifically designed for audio language modeling. It innovatively constructs a wide vector-quantized (VQ) codebook space, integrated with an extended-context Transformer, multi-scale GAN discriminators, and an inverse short-time Fourier transform (iSTFT)-based reconstruction architecture, enabling compression of 1-second 24 kHz audio into only 40–75 semantically rich tokens. WavTokenizer achieves state-of-the-art performance across speech, music, and general audio reconstruction: it attains a new industry-leading UTMOS score (+0.32), improves VQ codebook utilization by 37%, and significantly enhances both perceptual quality and semantic consistency. The model is lightweight, open-source, and natively compatible with downstream audio generation tasks.
本文提出了一种新的通信接口,通过任务相关的令牌控制比特生成和保护,以解决AI模型与通信系统之间的传输问题。
Existing acceleration methods for diffusion models often compromise classification capability when compressing computation, struggling to balance generation quality and discriminative performance. This work proposes BiGain—a training-free, plug-and-play framework that jointly enhances both generative and classification performance in accelerated diffusion for the first time. At its core lies a spectrum-aware dual-operator compression mechanism: Laplacian gating enables token merging, while interpolation-extrapolation-based KV downsampling preserves high-frequency details and maintains low- and mid-frequency semantics. On ImageNet-1K, BiGain achieves a 7.15% absolute improvement in classification accuracy and a 0.34 reduction in FID (a 1.85% relative improvement) under a 70% token compression rate, significantly advancing the trade-off between speed and accuracy.
This work addresses the unclear trade-offs among compression efficiency, structural inductive bias, and cross-domain robustness in large language model tokenizers. Viewing tokenization through an information-theoretic lens as structured compression, the authors propose a variant of Byte Pair Encoding (BPE) integrated with principles from compressed sensing and introduce metrics such as channel capacity utilization. They systematically analyze how vocabulary size and training data volume influence text entropy distribution and contextual predictability. Experimental results reveal that while increasing training data enhances token diversity, it simultaneously strengthens contextual predictability. The proposed framework effectively quantifies tokenizer performance, offering both theoretical grounding and practical guidance for designing general-purpose, compression-oriented tokenization strategies and downstream modeling.
This study addresses the computational inefficiency in time series language models arising from the unified treatment of time series and prompt tokens, which exhibit fundamentally different information structures. The work reveals an asymmetry in token importance: time series tokens contribute unevenly across the frequency spectrum, while the influence of prompt tokens diminishes with model depth. Leveraging this insight, the authors propose an adaptive token compression framework that dynamically compresses time series tokens via spectral analysis and progressively prunes prompt tokens in deeper layers, enabling hierarchical, frequency-aware, non-uniform budget allocation. Evaluated across forecasting, classification, imputation, and anomaly detection tasks, the method achieves up to 7.68× speedup and improves performance in 78% of experimental settings.
Existing neural audio codecs struggle to simultaneously preserve acoustic fidelity and capture semantic content, and incorporating semantic information often degrades reconstruction quality. To address this challenge, this work proposes STACodec, a unified codec framework that introduces a Semantic Token Allocation (STA) mechanism to inject semantic representations from self-supervised models into the first layer of Residual Vector Quantization (RVQ). Furthermore, a Semantic Pre-Distillation (SPD) module is designed to directly predict semantic tokens during inference without relying on external semantic tokenizers. This approach achieves a synergistic optimization of acoustic detail and semantic capability, maintaining high reconstruction fidelity while significantly enhancing performance on downstream semantic tasks.