NeuroLex: A Lightweight Domain Language Model for EEG Report Understanding and Generation

📅 2025-11-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
General-purpose language models struggle to capture the domain-specific syntactic structures, diagnostic reasoning patterns, and negation expressions inherent in clinical electroencephalography (EEG) reports. To address this, we propose NeuroLex—a lightweight, domain-adapted language model that jointly models span-level linguistic features and clinical reasoning patterns unique to EEG text. NeuroLex is built via span-masking pretraining on raw EEG reports followed by instruction tuning, yielding an EEG-aware language backbone. Experiments demonstrate that, at comparable parameter count, NeuroLex significantly outperforms general-purpose and biomedical baselines: it reduces perplexity by 12.3%, improves terminology-focused question answering accuracy by 18.7%, enhances robustness to negation, and achieves 2.1× higher label efficiency. Furthermore, NeuroLex supports clinical applications including report refinement, paragraph summarization, and interpretable neural decoding—establishing a novel paradigm for clinical EEG text understanding.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyMachine Learning: Neuro-Symbolic LearningCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Clinical electroencephalogram (EEG) reports encode domain-specific linguistic conventions that general-purpose language models (LMs) fail to capture. We introduce NeuroLex, a lightweight domain-adaptive language model trained purely on EEG report text from the Harvard Electroencephalography Database. Unlike existing biomedical LMs, NeuroLex is tailored to the linguistic and diagnostic characteristics of EEG reporting, enabling it to serve as both an independent textual model and a decoder backbone for multimodal EEG-language systems. Using span-corruption pretraining and instruction-style fine-tuning on report polishing, paragraph summarization, and terminology question answering, NeuroLex learns the syntax and reasoning patterns characteristic of EEG interpretation. Comprehensive evaluations show that it achieves lower perplexity, higher extraction and summarization accuracy, better label efficiency, and improved robustness to negation and factual hallucination compared with general models of the same scale. With an EEG-aware linguistic backbone, NeuroLex bridges biomedical text modeling and brain-computer interface applications, offering a foundation for interpretable and language-driven neural decoding.
Problem

Research questions and friction points this paper is trying to address.

General language models fail to capture EEG report domain-specific linguistic conventions
Existing biomedical models lack tailoring to EEG reporting linguistic and diagnostic characteristics
Current systems struggle with EEG interpretation syntax, reasoning patterns, and factual accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lightweight domain-adaptive language model for EEG
Trained using span-corruption and instruction fine-tuning
Specialized in EEG report understanding and generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kang Yin
Dept. of Artificial Intelligence, Korea University, Seoul, Republic of Korea
Hye-Bin Shin
Hye-Bin Shin
Korea University