Score
Designs, implements, and evaluates deterministic and stochastic transformations of audio signals and paired audio–text examples to increase training-data diversity, robustness, or to produce contrastive/negative examples for learning and decoding. Techniques include temporal, spectral, amplitude, speed and pitch perturbations, additive noise, structured or simulated acoustic mixtures, mixup of audio and audio–text pairs, and generation of targeted contrastive perturbations or negative branches for contrastive decoding.
This work addresses the susceptibility of large audio language models to hallucinations caused by linguistic priors overpowering acoustic evidence. To mitigate this, the authors propose a task- and sample-adaptive perturbation selection mechanism within a contrastive decoding framework. Leveraging a structured audio perturbation bank spanning temporal, spectral, frequency, and amplitude domains, the method dynamically selects optimal negative-sample perturbation strategies and employs a lightweight selector for efficient routing. The approach yields a 4.3% absolute improvement in accuracy on existence tasks and significantly boosts performance on temporal tasks from 74.7% to 81.4%. Furthermore, the study validates the efficacy of binary-constrained prompts, underscoring the critical role of adaptive perturbation strategies in alleviating hallucinations in audio language models.
Audio representation learning heavily relies on large-scale real-world recordings, while manual annotation and data augmentation struggle to capture the full diversity of physical acoustics. Method: We propose a synthetic-driven contrastive learning framework that requires no real audio data. It employs differentiable and stochastic sound synthesizers to generate physically consistent synthetic “twin” positive pairs online via causally interpretable parameter perturbations (e.g., timbre, pitch, envelope), thereby constructing high-diversity contrastive tasks. Contribution/Results: We introduce the first positive-pair construction paradigm grounded in causal perturbations of synthesizer parameters; require only a single interpretable hyperparameter and zero real-data storage; and achieve, for the first time, synthetic-data-only models that surpass real-data baselines on ESC-50, UrbanSound8K, and SpeechCommands. This significantly reduces data dependency and storage overhead, establishing a new paradigm for low-resource audio representation learning.
To address the challenges of subjective quality assessment and the inefficiency of purely data-driven models in speech/audio coding, this paper proposes a tightly integrated hybrid neural coding framework that synergistically combines model-driven and data-driven paradigms. Methodologically, it introduces a novel multi-level hybrid architecture that deeply couples psychoacoustic-weighted loss, customized time-frequency domain prediction (TF-Codec/MDCTNet), an LPCNet-based backbone, and a neural post-processing module, trained end-to-end via an autoencoder paradigm. The core contribution lies in systematically bridging the performance gap between classical signal modeling and end-to-end deep learning. Experimental results demonstrate that, at ultra-low bitrates of 1.6–3.2 kbps, the proposed method achieves a P.808 MOS gain of ≥0.5 over baselines, yielding subjective audio quality approaching that of wideband codecs, while increasing computational overhead by less than 15%.
Conventional self-supervised audio representation learning relies solely on clip-level sampling, leading to insufficient frame-level modeling capability. Method: This paper proposes a multi-granularity contrastive learning framework that jointly leverages clip-level, frame-level, and task-guided sampling to construct multi-perspective contrastive losses, enabling collaborative optimization of general-purpose audio representations. Contribution/Results: To our knowledge, this is the first work to incorporate both frame-level and task-specific sampling into self-supervised pre-training, overcoming the limitations of single-granularity representation learning. Pre-trained on a subset of AudioSet and evaluated via frozen-feature transfer to downstream tasks, our method achieves 25%, 20%, and 3.6% absolute improvements in clip classification, sound event detection, and pitch detection, respectively—demonstrating significantly enhanced fine-grained frame-level perception.
Neural speech enhancement models often suffer from poor generalization in real-world far-field scenarios. To address this, this paper proposes a near-field guided pseudo-label training paradigm. Leveraging real-recorded near-field–far-field speech pairs, it first trains a high-fidelity speech enhancement model on near-field data to generate high-quality pseudo-clean labels for corresponding far-field mixtures. These pseudo-labels then supervise the end-to-end training of a far-field enhancement model. Crucially, the approach eliminates the need for ground-truth far-field clean speech, enabling—for the first time—a purely real-data-driven far-field speech enhancement training framework. By bypassing synthetic data, it effectively bridges the distributional gap between simulated and real acoustic domains. Experiments on the CHiME-4 real-world dataset demonstrate that the generated pseudo-labels achieve high fidelity, and the resulting far-field model significantly outperforms conventional simulation-supervised baselines in terms of objective and perceptual metrics.
Although contrastive decoding (CD) has demonstrated performance improvements in large audio language models (LALMs), its underlying mechanisms and conditions for effectiveness remain unclear. This work systematically evaluates four CD strategies across diverse LALM architectures and introduces a transition-matrix-based framework to analyze error patterns. The analysis reveals that CD effectively corrects only two error types—“false negatives” (misclassifying audio as absent) and “uncertain guesses”—but fails to address erroneous reasoning or confidently incorrect predictions. Among the evaluated approaches, Audio-Aware Decoding and Audio Contrastive Decoding emerge as the most effective strategies. Building on these findings, the study establishes principled guidelines for selecting CD methods based on the baseline model’s characteristic error profiles, offering both theoretical insights and practical recommendations for optimizing LALM inference.
This work addresses the susceptibility of unified audio-language models to temporal smoothing bias during generation, which hinders their effective utilization of transient acoustic cues and results in insufficient fine-grained alignment between output text and audio. To mitigate this issue, the authors propose a training-free temporal contrastive decoding method that, at inference time, constructs a contrastive signal between the original input and a temporally blurred “slow-path” view to dynamically refine the logits of the next token. The approach introduces, for the first time, a self-normalized stability score coupled with an uncertainty-aware gating mechanism, integrating waveform-smoothed recoding, adaptive blurring windows, and token-level logit updates to enable on-demand, precise enhancement of transient audio information. Evaluated on the MMAU and AIR-Bench benchmarks, the method consistently improves performance across multiple strong baseline models, demonstrating both effectiveness and architectural generality.
Existing research on audio spoofing detection often overlooks real-world scenarios where speech coexists with environmental sounds and may be partially manipulated. To address this gap, this work introduces PC-Mix, the first dataset specifically designed for partial-component spoofing detection in mixed audio, featuring realistic scenarios with locally forged environmental sounds. We further propose a joint learning framework that operates across both speech and environmental sound components. Through a unified evaluation protocol, our experiments demonstrate that spoofing detection under mixed conditions is significantly more challenging, and models trained under target-matched conditions substantially outperform those directly transferred from single-component settings. This study fills a critical research void in partial spoofing detection involving environmental sounds and mixed acoustic conditions.
This study systematically compares generative and discriminative deep learning approaches for speech enhancement across varying signal-to-noise ratios, training data match conditions, and dataset scales. Through comprehensive evaluation of denoising performance, convergence speed, computational complexity, and speech hallucination—quantified via word error rate and phoneme similarity—the work reveals, for the first time, a multidimensional trade-off between robustness, efficiency, and perceptual quality. Generative models demonstrate superior perceptual quality but incur higher computational costs and greater hallucination risk compared to their discriminative counterparts. These findings provide empirical evidence and practical guidelines for selecting appropriate methods in real-world deployment scenarios.
This work addresses the challenge of detecting and localizing target sounds in complex acoustic scenes by proposing a unified encoder-based shared representation learning framework. Departing from conventional conditional embedding mechanisms, the method jointly encodes reference and mixture audio signals within a shared semantic space and employs multi-task learning to simultaneously optimize detection and localization performance. By innovatively aligning audio embeddings and enabling end-to-end training, the approach significantly enhances generalization to unseen sound categories while simplifying model architecture. Evaluated on the URBAN-SED dataset, the proposed method achieves a segment-level F1 score of 83.15% and an overall accuracy of 95.17%, establishing a new state-of-the-art performance.