continuous integrate-and-fire

Design and implement continuous integrate-and-fire (CIF) modules that learn monotonic soft alignments between high-rate input frame sequences (commonly speech features) and lower-rate token sequences by integrating frame-level weights and emitting token embeddings when an accumulator crosses a threshold. Build or analyze CIF-based downsampling layers and alignment extractors that reduce encoder time resolution while preserving frame-to-token alignment for downstream decoding (e.g., ASR) or token-level modeling.

continuousintegrate-and-fire

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR

Oct 25, 2025
RM
Ruixiang Mao
🏛️ Northeastern University | NiuTrans Research | Kunming University of Science and Technology

To address unstable cross-lingual (particularly English–French) acoustic-text alignment in non-autoregressive speech recognition using the continuous Integrate-and-Fire (CIF) mechanism, this paper proposes a Multi-scale CIF (M-CIF) model. M-CIF introduces progressive supervision signals at both phoneme and character levels within the continuous integration-and-firing framework, enabling multi-granularity alignment; it further enhances subword representation alignment robustness via knowledge distillation. Experiments on CommonVoice show that M-CIF significantly improves performance, reducing WER by 4.21% for German and 3.05% for French. Moreover, Peak Error (PE) and Segmentation Error (SE) metrics confirm the effectiveness of its hierarchical alignment modeling. To our knowledge, this is the first work to incorporate explicit multi-scale supervision into the CIF framework, establishing a more stable and interpretable alignment paradigm for cross-lingual non-autoregressive speech recognition.

Enhances acoustic-text alignment using multi-level phoneme supervisionImproves alignment stability for non-autoregressive speech recognitionReduces phonetic confusion and segmentation errors in multilingual ASR

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition

Jun 04, 2025
YX
Yinfeng Xia
🏛️ Honor Device Co, Ltd | Shanghai Jiao Tong University

To address the challenge of deploying large pre-trained speech models (e.g., Whisper) in low-latency streaming ASR, this paper proposes a prefix-to-prefix fine-tuning framework enabling quasi-monotonic speech-text alignment. Methodologically, it introduces: (1) a novel Continuous Integrate-and-Fire alignment mechanism; (2) Monotonic Finite Look-ahead Attention, enabling tunable latency–accuracy trade-offs; and (3) end-to-end streaming fine-tuning via wait-k decoding. Evaluated across multiple datasets, the approach achieves millisecond-level controllable latency while matching near-offline Whisper accuracy. Theoretically, we prove alignment monotonicity and training stability, establishing— for the first time—the first streaming fine-tuning paradigm for Whisper with strict, configurable latency guarantees.

Achieving controllable latency-quality trade-off in streaming applicationsEstablishing quasi-monotonic alignment between speech and text tokensIntegrating large pre-trained models into streaming speech recognition systems

This work addresses the challenge that speech large language models (Speech LLMs) struggle to accurately localize hotwords and long-tail named entities under weak supervision due to strong language model priors. To this end, we propose CLAR, a dual-encoder speech-text retrieval framework that, for the first time, leverages the Continuous Integrate-and-Fire (CIF) mechanism to achieve timestamp-free monotonic alignment at the token level in an unsupervised manner. CLAR further incorporates length-aware local matching to enhance acoustic cues for short entities. Through multi-granularity contrastive learning and CIF-based quantity constraints, our approach effectively mitigates representation dilution and attention drift. Experimental results demonstrate that CLAR significantly improves hotword retrieval accuracy and substantially reduces both character error rate (CER) and named entity word error rate (B-WER) over strong baselines.

contextual ASRhotword localizationnamed entity recognition

Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration

Nov 26, 2024
YH
Yuhang Han
🏛️ Northwestern Polytechnical University | Sichuan University | Westlake University | Zhejiang University

Multimodal large language models (MLLMs) suffer from quadratic computational and memory overhead due to long visual token sequences, severely hindering deployment efficiency. To address the limitations of existing training-free compression methods—particularly in redundant token identification and critical information recovery—we propose a decoupled three-stage “Filter–Compensate–Compress” (FCC) framework. First, redundant tokens are filtered based on visual feature similarity. Second, cross-token semantic correlations are modeled to adaptively compensate discarded information into retained tokens. Third, weighted attention-guided fusion mitigates semantic dilution during compression. The framework requires no fine-tuning or gradient updates and supports dual-path adaptation—FiCoCo-V for vision encoders and FiCoCo-L for language decoders. Evaluated on LLaVA-1.5-7B and NeXT-7B, FCC achieves up to 5.7× and 14.7× FLOPs reduction while preserving 92.8% and 93.6% of original performance, respectively—substantially outperforming state-of-the-art training-free approaches.

Identifies and recovers essential information from discarded tokens.Optimizes token reduction without requiring model retraining.Reduces computational and memory challenges in MLLMs.

Latest Papers

What's happening recently
View more

This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.

automatic speech recognitionencoder frame gridspeech LLMs

Diffusion-based large language models (dLLMs) suffer from high computational overhead during parallel decoding due to excessive redundant [MASK] tokens and repeated context. This work is the first to uncover the redundancy mechanism from the perspective of [MASK] tokens and proposes a position-preserving [MASK] compression method that significantly reduces computational cost while retaining structural information. Furthermore, it introduces terminal-aware context enhancement and context folding expansion techniques to naturally and efficiently support long contexts. Experiments on the LLaDA model series demonstrate that the proposed approach substantially accelerates decoding and improves generation quality with minimal additional computational overhead.

computational redundancycontext compressiondiffusion LLMs

This work addresses the instability and performance degradation in long-form speech recognition caused by the abrupt emergence of alignment information in deep layers of Aligner-Encoder models. To mitigate this issue, the authors propose InterAligner and InterCTC mechanisms that introduce progressive alignment objectives and CTC losses at intermediate encoder layers, enabling the first framework to learn alignment gradually along the depth dimension. Integrated into a 17-layer Conformer architecture, the joint optimization significantly enhances training stability and recognition accuracy, reducing word error rates on LibriSpeech from 5.0/7.8 to 3.1/5.6 on test-clean/test-other, with particularly pronounced improvements in long utterance scenarios.

Aligner-EncoderalignmentASR

Hot Scholars

TL

Tony Lindeberg

Professor of Computer Science - Computational Vision, KTH Royal Institute of Technology
Computer VisionScale SpaceRecognitionImage Analysis
JE

Jens Egholm Pedersen

KTH Royal Institute of Technology
neuromorphiccomputer scienceevent based visionmachine learning
MW

Maciej Wielgosz

AGH University of Science and Technology
Cognitive ComputingMachine LearningDeep LearningHardware Acceleration
TS

Tony Shardlow

Department of Mathematical Sciences, University of Bath
Numerical analysisstochastic differential equations
EN

Emre Neftci

Institute Director, Forschungszentrum Jülich; Professor, RWTH Aachen
Neuromorphic EngineeringComputational NeuroscienceCognitive Systems and BehaviorMachine Learning