blind source separation

Techniques for decomposing observed mixed signals into underlying latent source signals without complete knowledge of the mixing process; applied to denoise bioelectrical waveforms, separate interfering audio components after unknown propagation, or disentangle intertwined multimodal embeddings.

blindsourceseparation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Disentangling Modes and Interference in the Spectrogram of Multicomponent Signals

Mar 19, 2025
KP
K'evin Polisano
🏛️ Univ. Grenoble Alpes | Ens de Lyon | Université Caen Normandie

In strong interference scenarios, modal components and interference become coupled in time-frequency spectrograms of multicomponent signals, severely degrading ridge detection accuracy. Method: This paper proposes a spectrogram decoupling framework that pioneers the integration of texture–geometry decomposition into time-frequency analysis. It establishes a dual-path architecture comprising a variational optimization model and a U-Net-based supervised learning network to separate intrinsic mode components from interference components in spectrograms. Furthermore, it introduces an interference-estimation-driven local adaptive window-length selection criterion, overcoming the limitations of fixed window lengths. Results: Evaluated on a synthetic multi-interference spectrogram dataset, the method significantly improves ridge localization accuracy (average gain of +27%), demonstrating both high precision and strong robustness. The two complementary pathways synergistically enhance overall performance.

Decompose spectrogram into mode and interference partsEnhance time-frequency analysis under strong interferenceImprove ridge detection in spectrograms with close modes

Soft Disentanglement in Frequency Bands for Neural Audio Codecs

Oct 04, 2025
BG
Benoit Ginies
🏛️ LTCI | Telecom Paris | Institut polytechnique de Paris

Neural audio codecs often rely on data- or task-specific priors to disentangle frequency-band features, resulting in poor interpretability and limited generalizability. Method: We propose a generic soft disentanglement representation learning framework. It first applies spectral decomposition to project time-domain audio into orthogonal frequency-band subspaces; then employs a multi-branch encoder to model each band independently, jointly optimized via reconstruction and perceptual losses. Crucially, the framework imposes no assumptions about task structure or data distribution, enabling task-agnostic, soft intra-band semantic disentanglement. Contribution/Results: Experiments demonstrate significant improvements over state-of-the-art codecs in objective audio quality metrics (e.g., PESQ, STOI) and perceptual fidelity. Moreover, the learned representations exhibit strong generalization to downstream tasks—such as audio inpainting—and provide interpretable, structured frequency-band semantics without architectural or prior constraints.

Achieving disentangled representations in neural audio codecsEnhancing reconstruction and perceptual quality for audio tasksImproving model interpretability through frequency band decomposition

Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux

Oct 04, 2025
BG
Benoît Giniès
🏛️ Télécom Paris | Institut polytechnique de Paris

Weak interpretability of representations and dataset- and task-specific disentanglement remain critical bottlenecks in neural audio codecs. To address this, we propose a time-domain neural codec based on spectral decomposition of the input signal, incorporating a soft frequency-band decoupling mechanism that explicitly models semantic independence across frequency subbands during encoding. Our approach further integrates spectral prior structures to guide representation learning. A differentiable separation loss enables joint time–frequency modeling, substantially improving semantic disentanglement and cross-task generalization. Experiments demonstrate that our model surpasses state-of-the-art baselines in reconstruction fidelity (STOI, PESQ) and perceptual quality (MOS). It exhibits strong robustness and versatility across diverse downstream tasks—including speech enhancement, source separation, and audio synthesis—without task-specific architectural modifications.

Addressing dataset dependency in disentanglement techniquesEnhancing interpretability of neural audio codec representationsImproving reconstruction fidelity and perceptual quality

Blind Source Separation of Single-Channel Mixtures via Multi-Encoder Autoencoders

Aug 31, 2023
MB
Matthew B. Webster
🏛️ Mellowing Factory Co. Ltd.

Addressing the long-standing challenge of single-channel nonlinear blind source separation (BSS), this paper proposes a novel multi-encoder autoencoder framework—the first to introduce a multi-encoder architecture for this task. The method centers on three key components: (i) an encoding mask mechanism that enforces specialization of latent representations into distinct source-specific subspaces; (ii) a sparse mixing loss that regularizes the mixing structure via sparsity-inducing priors; and (iii) a zero-reconstruction loss that ensures orthogonality and interpretability of recovered sources. The model is trained end-to-end to independently estimate and separate constituent sources from their nonlinear mixture. Extensive evaluation on synthetic data and real-world multichannel polysomnography (PSG) recordings—specifically electrocardiogram (ECG) and photoplethysmogram (PPG)-derived respiratory signals—demonstrates substantial performance gains over conventional single-channel BSS approaches. This work establishes a new paradigm for physiological signal disentanglement.

Develop novel encoding masking and loss functions for source inferenceLeverage multi-encoder autoencoders for feature subspace specializationSeparate sources from single-channel non-linear mixtures blindly

Naturalistic Music Decoding from EEG Data via Latent Diffusion Models

May 15, 2024
EP
Emilian Postolache
🏛️ Ca' Foscari University of Venice | Sony CSL | Tokyo Institute of Technology | Sapienza University of Rome

This study addresses the challenge of end-to-end reconstruction of high-fidelity, polyphonic, harmonically rich natural music directly from raw, non-invasive EEG signals. Methodologically, it introduces latent diffusion models (LDMs) to the EEG-to-audio decoding task for the first time, enabling direct mapping from raw temporal EEG to complex audio waveforms—without handcrafted features, channel selection, or signal preprocessing. A novel neural-embedding-based evaluation metric is proposed to better quantify auditory perceptual consistency. Experiments on the NMED-T dataset demonstrate substantial improvements in timbral and structural fidelity; reconstructed audio exhibits superior intelligibility and musicality compared to conventional linear models and VAE baselines. Key contributions include: (1) the first LDM framework explicitly designed for natural-music synthesis from EEG; (2) a full-brain, end-to-end decoding paradigm that maps raw EEG to high-quality audio; and (3) a semantics-aware evaluation framework tailored for neural decoding.

Brain SignalsHigh-quality Complex SignalsNatural Music Reconstruction

Latest Papers

What's happening recently
View more

EEG-D3: A Solution to the Hidden Overfitting Problem of Deep Learning Models

Dec 15, 2025
SL
Siegfried Ludwig
🏛️ Imperial College London | Aristotle University of Thessaloniki | National and Kapodistrian University of Athens

Deep learning–based EEG decoding often suffers from latent overfitting, severely impairing cross-scenario generalization and hindering real-world BCI deployment. To address this, we propose the weakly supervised Disentangled Decoding Decomposition (D3) framework. D3 leverages positional prediction within trial sequences to drive latent-space disentanglement; introduces a novel independent subnetwork architecture coupled with a nonlinear ICA–inspired disentanglement mechanism to isolate task-agnostic neural components; and establishes a cross-dataset component activation contrast paradigm that explicitly models linear separability in the latent space. Experiments demonstrate that D3 stably isolates physiologically meaningful components in motor imagery data while suppressing task-irrelevant artifacts. In sleep staging, D3 achieves significant few-shot accuracy gains—requiring only minimal labeled data—and generalizes robustly to real-world scenarios.

Addresses hidden overfitting in EEG deep learning modelsEnables generalization with minimal labeled dataSeparates latent brain activity components from artifacts

This work proposes a knowledge-driven, self-supervised approach to audio segmentation and source separation that circumvents the reliance on large-scale manually annotated data. By integrating external prior knowledge—such as musical scores—into the audio processing pipeline for the first time, the method leverages hidden Markov models to achieve effective segmentation and separation of music and film audio without requiring labeled training data. Evaluated on synthetic datasets, the approach demonstrates strong performance, and in real-world film soundtrack tests, it significantly outperforms purely data-driven methods when incorporating sound-class priors. This represents a notable advance toward annotation-free audio analysis through principled integration of domain knowledge.

audio source separationcinematic audioknowledge-driven

Existing research on audio spoofing detection often overlooks real-world scenarios where speech coexists with environmental sounds and may be partially manipulated. To address this gap, this work introduces PC-Mix, the first dataset specifically designed for partial-component spoofing detection in mixed audio, featuring realistic scenarios with locally forged environmental sounds. We further propose a joint learning framework that operates across both speech and environmental sound components. Through a unified evaluation protocol, our experiments demonstrate that spoofing detection under mixed conditions is significantly more challenging, and models trained under target-matched conditions substantially outperform those directly transferred from single-component settings. This study fills a critical research void in partial spoofing detection involving environmental sounds and mixed acoustic conditions.

environmental soundmixed audiopartial audio spoofing

This work addresses the challenge of learning disentangled and interpretable latent representations from complex, non-stationary, high-dimensional time-varying signals, which exhibit rich time-frequency structures that conventional variational autoencoders (VAEs) struggle to model effectively. To this end, we propose the Decompositional Variational Autoencoder (DecVAE), which uniquely integrates signal decomposition priors directly into the VAE framework. DecVAE employs an encoder-only architecture and jointly leverages signal decomposition models, contrastive self-supervised tasks, and variational inference to learn multi-subspace latent representations aligned with the intrinsic time-frequency characteristics of the data. Extensive experiments on synthetic data and three scientific datasets demonstrate that DecVAE substantially outperforms existing VAE approaches, achieving significant improvements in disentanglement quality, cross-task generalization, and interpretability of the learned latent representations.

disentangled representationhigh-dimensional datalatent generative mechanisms

This work addresses the misalignment between disentangled representations in existing music generation models and their intended semantic meanings, which limits controllable generation. It presents the first multidimensional disentanglement evaluation framework for music audio, assessing representations along four axes—informativeness, equivariance, invariance, and disentanglement—through probing analyses and a range of unsupervised disentanglement strategies, including inductive biases, data augmentation, adversarial objectives, and staged training. The study systematically evaluates how well key attributes such as musical structure and timbre are captured in learned representations. Experimental results reveal significant discrepancies between the actual behavior of latent embeddings and their design intent, highlighting substantial challenges in achieving semantic consistency and offering new directions for improving controllable music generation mechanisms.

controllable music generationdisentangled representationsembedding semantics

Hot Scholars

TV

Tuomas Virtanen

Tampere University
machine listeningaudio signal processingaudio
ZQ

Zhong-Qiu Wang

Associate Professor, Southern University of Science and Technology
Computer AuditionSpeech SeparationMicrophone ArrayAudio Signal Processing
YH

Yuan-Hao Wei

Hong Kong Polytechnic University
Machine LearningVAEICA
SG

Sharon Gannot

Prof. of Data Engineering, Bar-Ilan University, Israel
Acoustic signal processingAudio-video processingMachine learning for audio processing
GH

Gongping Huang

Professor, Wuhan University, Wuhan, China
Acoustic Signal ProcessingMicrophone ArraysSpeech EnhancementNoise Reduction