multi-branch acoustic modeling

Designs, builds, and evaluates multi-branch acoustic neural networks that process different acoustic feature streams in feature-specific branches to learn specialized representations and fuse branch outputs for final prediction. Includes selecting branch architectures, fusion strategies, and training techniques to achieve robust acoustic modeling across heterogeneous audio conditions.

multi-branchacousticmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of hierarchical consistency and fine-grained sub-label classification in heterogeneous audio within the Broad Sound Taxonomy (BST). To tackle these issues, the authors propose a unified framework that integrates multi-branch heterogeneous modeling, a hierarchy-aware classifier, and KNN-based post-processing. Leveraging CLAP audio-text representations alongside complementary acoustic features such as log-STFT, the method enhances fine-grained classification accuracy while preserving hierarchical structural constraints. Evaluated on BSD10k-v1.2, the single-model variant achieves a hierarchical F1 score of 80.84%, and an ensemble system further improves this to 81.25%, substantially outperforming existing approaches.

audio taxonomyBroad Sound TaxonomyDCASE challenge

Machine Learning in Acoustics: A Review and Open-Source Repository

Jul 06, 2025
RA
Ryan A. McCarthy
🏛️ Scripps Institution of Oceanography, UC San Diego | University of Rochester | Technical University of Denmark

This paper addresses the growing need for automated pattern recognition and modeling in acoustic data analysis. It systematically reviews state-of-the-art machine learning techniques—including deep learning, generative models, and physics-informed neural networks—for acoustic classification, regression, and generation tasks. To bridge the gap between methodology and application, we introduce AcousticsML, an open-source library featuring reproducible Jupyter notebooks and standardized preprocessing pipelines, with unified data interfaces and model evaluation protocols. AcousticsML enables end-to-end acoustic signal analysis and physics-constrained modeling, substantially lowering the barrier to entry for interdisciplinary researchers. By integrating data-driven and mechanism-driven paradigms, the library fosters an open, collaborative research framework for acoustics. It provides scalable, verifiable technical support for real-world applications including environmental noise monitoring, speech enhancement, and structural health diagnostics. (149 words)

Demonstrating Python-based ML techniques for acoustic dataProviding open-source examples to encourage reproducible approachesSurveying ML advances in acoustics for scientific insights

To address the challenges of subjective quality assessment and the inefficiency of purely data-driven models in speech/audio coding, this paper proposes a tightly integrated hybrid neural coding framework that synergistically combines model-driven and data-driven paradigms. Methodologically, it introduces a novel multi-level hybrid architecture that deeply couples psychoacoustic-weighted loss, customized time-frequency domain prediction (TF-Codec/MDCTNet), an LPCNet-based backbone, and a neural post-processing module, trained end-to-end via an autoencoder paradigm. The core contribution lies in systematically bridging the performance gap between classical signal modeling and end-to-end deep learning. Experimental results demonstrate that, at ultra-low bitrates of 1.6–3.2 kbps, the proposed method achieves a P.808 MOS gain of ≥0.5 over baselines, yielding subjective audio quality approaching that of wideband codecs, while increasing computational overhead by less than 15%.

Efficiency ImprovementNeural Voice and Audio CodingQuality Evaluation

This work addresses the performance gap between convolutional neural networks (CNNs) and pure Transformer architectures in end-to-end raw audio classification—specifically, the inferior accuracy of attention-only models lacking convolutional frontends. To bridge this gap, we propose three key innovations: (1) multi-scale temporal embeddings that jointly capture fine- and coarse-grained temporal structures; (2) a learnable, nonlinear, variable-bandwidth filterbank that replaces handcrafted STFT and fixed preprocessing pipelines; and (3) a CNN-inspired adaptive time-frequency pooling mechanism to enhance local invariance and improve feature compression. Evaluated on FreeSound 50K (200 classes) without pretraining, our model achieves new state-of-the-art performance, significantly outperforming leading CNNs and hybrid architectures in mean average precision (mAP). This is the first demonstration that a purely attention-based model can surpass conventional approaches on raw-audio understanding tasks—establishing both its superiority and practical feasibility.

Applying Transformer architectures to raw audio signals without convolutional layersImproving Transformer performance using convolutional-inspired pooling and wavelet-based multi-rate signal processingOutperforming convolutional models on audio classification tasks without unsupervised pre-training

Tailored Design of Audio-Visual Speech Recognition Models using Branchformers

Jul 09, 2024
DG
David Gimeno-G'omez
🏛️ Universitat Polit`ecnica de Val`encia

To address the high computational complexity and poor interpretability of cross-modal interactions in audio-visual speech recognition (AVSR) under noisy conditions, this paper introduces Branchformer—the first application of this architecture to AVSR—proposing a novel two-stage, customized unified encoder-decoder framework: “unimodal-first, then fusion.” By incorporating modality-specific branch scoring and layer-level structural pruning, the method achieves parameter-efficient and interpretable audio-visual joint modeling. Evaluated on multi-scenario English and Spanish benchmarks, it attains word error rates (WER) of 2.5% and 9.1%, respectively—significantly outperforming comparable large models while reducing parameter count substantially and achieving state-of-the-art performance. Key contributions include: (1) the pioneering adaptation of Branchformer to AVSR; (2) a new architectural paradigm that jointly optimizes model lightweighting and cross-modal interpretability; and (3) enhanced end-to-end cross-modal collaborative modeling capability.

Achieve state-of-the-art recognition rates efficiently.Optimize cross-modal architecture for AVSR.Reduce model complexity and computational cost.

Latest Papers

What's happening recently
View more

Acoustic neural networks: Identifying design principles and exploring physical feasibility

Nov 26, 2025
IK
Ivan Kalthoff
🏛️ RWTH Aachen University | University of Münster

Current acoustic neural networks lack a systematic design framework that explicitly links learnable parameters to physically measurable acoustic properties—such as material attenuation and geometric configuration—hindering the deployment of low-power, passive acoustic computing in resource-constrained environments. To address this, we propose the first physics-aware digital twin training framework that incorporates hardware-imposed acoustic constraints—including non-negativity and zero bias—directly into network optimization, while establishing an explicit mapping from network weights to measurable acoustic parameters. Furthermore, we introduce the SincHSRNN, a hybrid model compatible with passive components, integrating acoustic waveguide modeling, intensity-based nonlinearity, learnable bandpass filtering, and hierarchical temporal processing. Evaluated on AudioMNIST, it achieves 95% accuracy—the first demonstration of efficient speech recognition using purely passive acoustic hardware—thereby unifying physical realizability with competitive computational performance.

Developing systematic design framework for acoustic neural networksEnabling low-power computation with wave-based analog systemsEstablishing physical feasibility through measurable acoustic properties

This work addresses the high computational complexity of existing artificial neural network (ANN)-based methods and the information loss inherent in spiking neural network (SNN)-based approaches due to binary activation in single-channel speech enhancement. To overcome these limitations, the authors propose a hybrid ANN/SNN dual-branch architecture that leverages BandSplit for subband decomposition and TF-Mamba for modeling time-frequency dependencies. The design incorporates a Spiking Feature Extraction Group (SFEG) and an Information Transformation Block (ITB), along with a time-frequency adaptive cross-attention fusion mechanism to enable efficient collaboration between the two branches. Evaluated on three public datasets, the proposed model achieves superior speech enhancement performance while reducing average computational complexity by 7.5× compared to baseline models, substantially improving energy efficiency.

computational complexityenergy consumptionmodel performance

This work proposes an end-to-end, feature-free audio classification approach based on a parallel deep reservoir computing architecture that operates directly on raw audio waveforms, eliminating the need for explicit feature extraction such as MFCCs. Traditional methods relying on handcrafted features often incur high computational overhead and complex preprocessing pipelines. To evaluate the efficacy of the proposed design, the authors conduct comparative experiments using shallow, serial, and parallel deep reservoir models. Results demonstrate that the parallel architecture achieves significantly superior performance over baseline methods while maintaining low model complexity. The approach enables efficient temporal modeling and hierarchical representation learning, highlighting its scalability and practical potential for audio processing tasks.

acoustic signal preprocessingend-to-end classificationfeature-free

This work proposes GPA, a general-purpose audio model that unifies speech synthesis, automatic speech recognition, and voice conversion within a single autoregressive Transformer architecture—addressing the fragmentation, poor scalability, and limited generalization of traditional task-specific speech systems. By leveraging a shared discrete speech token space and an instruction-driven mechanism, GPA enables zero-architecture-modification task switching. The model employs multi-task joint training and a high-throughput inference pipeline, facilitating lightweight deployment. Experimental results demonstrate that GPA achieves competitive performance across multiple tasks, with its 0.3B-parameter variant particularly well-suited for low-latency, resource-constrained edge scenarios.

autoregressive transformersspeech recognitionspeech synthesis

This study addresses the challenge of balancing accuracy and model compactness in underwater acoustic classification, alongside concerns regarding generalization due to insufficiently rigorous evaluation protocols. To this end, we propose a compact and efficient classification framework that integrates auditory-inspired time-frequency and cochlear multi-representation feature engineering, temporal statistical pooling, and a lightweight convolutional architecture. Furthermore, a strict recording-level data partitioning protocol is introduced to ensure reliable evaluation. Experimental results demonstrate that the proposed framework achieves an F1-score of 0.9918 on the ShipsEar dataset. On the DeepShip dataset, a small model with only 157K parameters attains an F1-score of 0.7226, outperforming larger counterparts and thereby validating its effectiveness for efficient deployment under stringent evaluation conditions.

Acoustic ClassificationCompact ModelsDeployability

Hot Scholars

YT

Yu Tsao

Research Fellow (Professor), Deputy Director, CITI, Academia Sinica
Assistive Oral Communication TechnologiesSpeech EnhancementVoice ConversionSpeech Assessment
AW

Alexander Waibel

Carnegie Mellon, KIT, Karlsruhe Institute of Technology, University of Karlsruhe
Machine LearningNeural NetworksSpeech TranslationMultimodal Interfaces
JN

Jan Niehues

Institute for Anthropomatics and Robotics (IAR), Karlsruhe Institute for Technology (KIT)
natural language processing - machine translation
NM

Nobuaki Minematsu

The University of Tokyo
Speech CommunicationForeign Language Learning
RH

Reinhold Haeb-Umbach

Professor of Communications Engineering, University of Paderborn
automatic speech recognitionspeech enhancementstatistical signal processing