Score
Design and implement continuous integrate-and-fire (CIF) modules that learn monotonic soft alignments between high-rate input frame sequences (commonly speech features) and lower-rate token sequences by integrating frame-level weights and emitting token embeddings when an accumulator crosses a threshold. Build or analyze CIF-based downsampling layers and alignment extractors that reduce encoder time resolution while preserving frame-to-token alignment for downstream decoding (e.g., ASR) or token-level modeling.
To address unstable cross-lingual (particularly English–French) acoustic-text alignment in non-autoregressive speech recognition using the continuous Integrate-and-Fire (CIF) mechanism, this paper proposes a Multi-scale CIF (M-CIF) model. M-CIF introduces progressive supervision signals at both phoneme and character levels within the continuous integration-and-firing framework, enabling multi-granularity alignment; it further enhances subword representation alignment robustness via knowledge distillation. Experiments on CommonVoice show that M-CIF significantly improves performance, reducing WER by 4.21% for German and 3.05% for French. Moreover, Peak Error (PE) and Segmentation Error (SE) metrics confirm the effectiveness of its hierarchical alignment modeling. To our knowledge, this is the first work to incorporate explicit multi-scale supervision into the CIF framework, establishing a more stable and interpretable alignment paradigm for cross-lingual non-autoregressive speech recognition.
To address the challenge of deploying large pre-trained speech models (e.g., Whisper) in low-latency streaming ASR, this paper proposes a prefix-to-prefix fine-tuning framework enabling quasi-monotonic speech-text alignment. Methodologically, it introduces: (1) a novel Continuous Integrate-and-Fire alignment mechanism; (2) Monotonic Finite Look-ahead Attention, enabling tunable latency–accuracy trade-offs; and (3) end-to-end streaming fine-tuning via wait-k decoding. Evaluated across multiple datasets, the approach achieves millisecond-level controllable latency while matching near-offline Whisper accuracy. Theoretically, we prove alignment monotonicity and training stability, establishing— for the first time—the first streaming fine-tuning paradigm for Whisper with strict, configurable latency guarantees.
This work addresses the challenge that speech large language models (Speech LLMs) struggle to accurately localize hotwords and long-tail named entities under weak supervision due to strong language model priors. To this end, we propose CLAR, a dual-encoder speech-text retrieval framework that, for the first time, leverages the Continuous Integrate-and-Fire (CIF) mechanism to achieve timestamp-free monotonic alignment at the token level in an unsupervised manner. CLAR further incorporates length-aware local matching to enhance acoustic cues for short entities. Through multi-granularity contrastive learning and CIF-based quantity constraints, our approach effectively mitigates representation dilution and attention drift. Experimental results demonstrate that CLAR significantly improves hotword retrieval accuracy and substantially reduces both character error rate (CER) and named entity word error rate (B-WER) over strong baselines.
Multimodal large language models (MLLMs) suffer from quadratic computational and memory overhead due to long visual token sequences, severely hindering deployment efficiency. To address the limitations of existing training-free compression methods—particularly in redundant token identification and critical information recovery—we propose a decoupled three-stage “Filter–Compensate–Compress” (FCC) framework. First, redundant tokens are filtered based on visual feature similarity. Second, cross-token semantic correlations are modeled to adaptively compensate discarded information into retained tokens. Third, weighted attention-guided fusion mitigates semantic dilution during compression. The framework requires no fine-tuning or gradient updates and supports dual-path adaptation—FiCoCo-V for vision encoders and FiCoCo-L for language decoders. Evaluated on LLaVA-1.5-7B and NeXT-7B, FCC achieves up to 5.7× and 14.7× FLOPs reduction while preserving 92.8% and 93.6% of original performance, respectively—substantially outperforming state-of-the-art training-free approaches.
This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.
Diffusion-based large language models (dLLMs) suffer from high computational overhead during parallel decoding due to excessive redundant [MASK] tokens and repeated context. This work is the first to uncover the redundancy mechanism from the perspective of [MASK] tokens and proposes a position-preserving [MASK] compression method that significantly reduces computational cost while retaining structural information. Furthermore, it introduces terminal-aware context enhancement and context folding expansion techniques to naturally and efficiently support long contexts. Experiments on the LLaDA model series demonstrate that the proposed approach substantially accelerates decoding and improves generation quality with minimal additional computational overhead.
This work addresses the instability and performance degradation in long-form speech recognition caused by the abrupt emergence of alignment information in deep layers of Aligner-Encoder models. To mitigate this issue, the authors propose InterAligner and InterCTC mechanisms that introduce progressive alignment objectives and CTC losses at intermediate encoder layers, enabling the first framework to learn alignment gradually along the depth dimension. Integrated into a 17-layer Conformer architecture, the joint optimization significantly enhances training stability and recognition accuracy, reducing word error rates on LibriSpeech from 5.0/7.8 to 3.1/5.6 on test-clean/test-other, with particularly pronounced improvements in long utterance scenarios.