🤖 AI Summary
This study addresses the limitations of existing intracranial EEG (iEEG) foundation models, which often neglect explicit physiological features and exhibit insufficient clinical generalizability. To this end, we propose PHASE, a framework that employs a hierarchical architecture with a masked latent prediction mechanism to incorporate temporal and spatiotemporal physiological features as explicit learning objectives during pretraining. This approach effectively captures critical neural information without requiring epilepsy or anatomical labels. Evaluated on the Omni-iEEG benchmark, PHASE comprehensively outperforms existing models and establishes new state-of-the-art results across multiple tasks. Notably, its frozen representations yield substantially improved performance, demonstrating superior cross-institutional zero-shot generalization capabilities.
📝 Abstract
Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31\%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.