Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG

📅 2026-07-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of modeling multiscale neural signals and achieving interpretable sleep staging in real-world sparse electroencephalography (EEG) data by proposing RG-Flow Transformer, which incorporates renormalization group (RG) inductive bias into the Transformer architecture for the first time. The model leverages learnable anomalous dimensions, block-spin coarse-graining, and an entropy-gated synchronization bridge to effectively capture the scale invariance inherent in EEG signals. Evaluated on the Sleep-EDF dataset, RG-Flow achieves comparable five-class sleep staging accuracy to the standard Transformer (77.3% vs. 77.0%) while significantly recovering the continuous spectral exponent β (R² = 0.416), thereby offering enhanced interpretability that the conventional Transformer lacks.
📝 Abstract
Brain field potentials are scale-free: their power spectra follow a $1/f^β$ law whose aperiodic exponent $β$ tracks cortical state, and sleep depth in particular is a shift in $β$. We ask whether a transformer endowed with an explicit renormalization-group (RG) inductive bias -- the RG-Flow Transformer, which couples ordinary self-attention to a scale-aware stream with a learnable anomalous dimension $γ$, block-spin coarse-graining, and an entropy-gated synchronization bridge -- has an advantage over a parameter-matched vanilla transformer on \emph{real, scarce} EEG. Using the PhysioNet Sleep-EDF corpus with a strict leakage-free by-subject hold-out, we (i) benchmark RG-Flow against a param-matched vanilla transformer and a hierarchy-only ablation on 5-class AASM sleep staging, (ii) sweep the per-subject data budget to look for the inductive-bias crossover predicted when data are scarce, and (iii) test whether RG-Flow's learned $γ$ tracks the measured spectral exponent $β$ out-of-sample -- a quantity the vanilla model does not possess. Across $5$ subjects and $5$ seeds under leave-one-subject-out cross-validation, RG-Flow and the vanilla transformer are statistically indistinguishable on 5-class staging (77.3\% vs 77.0\% accuracy; paired $p=0.294$), and the predicted scarce-data crossover does not appear: vanilla is numerically ahead at every data-limited budget. What does separate the models is interpretability -- RG-Flow recovers the continuous spectral exponent out-of-sample ($β$-recovery $R^2 = 0.416$), a capability the vanilla architecture has no analogue for.
Problem

Research questions and friction points this paper is trying to address.

scale-free EEG
sleep staging
spectral exponent
scarce neural data
interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

RG-Flow Transformer
scale-aware attention
renormalization group
spectral exponent
scarce neural data