NeuroClean: A Generalized Machine-Learning Approach to Neural Time-Series Conditioning

📅 2025-11-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
EEG/LFP signals are often severely contaminated by artifacts and noise, and conventional preprocessing relies heavily on manual intervention, compromising reproducibility. To address this, we propose a fully automated, unsupervised five-step pipeline: integrated bandpass and notch filtering, automatic bad-channel detection, ICA decomposition, and clustering-based automatic component classification. Our key contribution lies in the deep integration of ICA with unsupervised machine learning to achieve high-accuracy artifact identification and removal while preserving task-relevant neural information—thereby substantially reducing human bias. Evaluated across multiple heterogeneous datasets, our method improves classification accuracy of downstream logistic regression models—from 74% with conventionally preprocessed data to over 97%—demonstrating marked gains in performance and reliability for brain–computer interface and neural decoding applications.

Technology Category

Cognitive Modeling & Cognitive Systems: Neural Spike CodingMachine Learning: Semi-Supervised LearningHumans and AI: Brain-Sensing and Analysis

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsResponsible Web: Machine-in-the-loop, human agency and autonomyWeb Mining and Content Analysis: Web data integration and cleaning
📝 Abstract
Electroencephalography (EEG) and local field potentials (LFP) are two widely used techniques to record electrical activity from the brain. These signals are used in both the clinical and research domains for multiple applications. However, most brain data recordings suffer from a myriad of artifacts and noise sources other than the brain itself. Thus, a major requirement for their use is proper and, given current volumes of data, a fully automatized conditioning. As a means to this end, here we introduce an unsupervised, multipurpose EEG/LFP preprocessing method, the NeuroClean pipeline. In addition to its completeness and reliability, NeuroClean is an unsupervised series of algorithms intended to mitigate reproducibility issues and biases caused by human intervention. The pipeline is designed as a five-step process, including the common bandpass and line noise filtering, and bad channel rejection. However, it incorporates an efficient independent component analysis with an automatic component rejection based on a clustering algorithm. This machine learning classifier is used to ensure that task-relevant information is preserved after each step of the cleaning process. We used several data sets to validate the pipeline. NeuroClean removed several common types of artifacts from the signal. Moreover, in the context of motor tasks of varying complexity, it yielded more than 97% accuracy (vs. a chance-level of 33.3%) in an optimized Multinomial Logistic Regression model after cleaning the data, compared to the raw data, which performed at 74% accuracy. These results show that NeuroClean is a promising pipeline and workflow that can be applied to future work and studies to achieve better generalization and performance on machine learning pipelines.
Problem

Research questions and friction points this paper is trying to address.

Automated conditioning of EEG and LFP signals to remove artifacts and noise
Unsupervised preprocessing to mitigate human intervention biases and reproducibility issues
Preserving task-relevant information while cleaning neural time-series data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised preprocessing pipeline for EEG/LFP signals
Automatic artifact rejection using clustering algorithms
Machine learning classifier preserves task-relevant information
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Manuel A. Hernandez Alonso
Departament de Matemàtiques i Informàtica, Universitat de Barcelona, Gran Via de les Corts Catalanes 585, 08007 Barcelona, Catalonia, Spain
M
M. Depass
Departament de Matemàtiques i Informàtica, Universitat de Barcelona, Gran Via de les Corts Catalanes 585, 08007 Barcelona, Catalonia, Spain
S
S. Quessy
Department of Neuroscience, Chemin de la Tour, Montreal, QC H3T 1J4, Canada
N
N. Dancause
Department of Neuroscience, Chemin de la Tour, Montreal, QC H3T 1J4, Canada
I
Ignasi Cos
Departament de Matemàtiques i Informàtica, Universitat de Barcelona, Gran Via de les Corts Catalanes 585, 08007 Barcelona, Catalonia, Spain