SPD-MetaFormer is what you need for small-data brain decoding

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high noise, significant inter-subject variability, and unclear attention contributions in few-shot brain signal decoding by proposing SPD-MetaFormer. Revealing that existing manifold attention weights tend toward uniform distributions, this work pioneers an attention-free architecture for symmetric positive definite (SPD) matrices, replacing complex adaptive learning with fixed geometric aggregation. Grounded in Log-Euclidean geometry, the method integrates geodesic residual updates, shared spectral feedforward mappings, and tangent space classification, processing covariance matrices through uniformly weighted Fréchet aggregation to substantially simplify the model for limited data regimes. Experiments across three EEG benchmarks demonstrate that this architecture achieves performance comparable to mainstream baselines, validating the effectiveness of minimalist SPD models in few-shot scenarios.
📝 Abstract
Brain signal decoding is challenging because neural recordings are noisy and vary across individuals, while labeled data are often limited. Recent attention-based models on the symmetric positive definite (SPD) manifold have nevertheless achieved strong performance using covariance and connectivity representations, yet the contribution of learned token weighting remains unclear. We examine two representative architectures, MAtt (based on log-Euclidean geometry) and GBWAtt (based on generalized Bures--Wasserstein geometry), and find that their learned attention weights remain close to uniform after training. We relate this behavior to bounded similarity parameterizations that, under the original softmax scaling, limit attention-weight contrast. Moreover, replacing learned weights with uniform weights, throughout training and evaluation, has little effect on mean predictive performance while preserving each model's original aggregation geometry. Motivated by these findings, we introduce SPD-MetaFormer, an attention-free architecture built on uniformly weighted Fréchet aggregation under log-Euclidean geometry. Its backbone uses a geodesic residual to update a summary token and a shared spectral feedforward map to transform all tokens, followed by a learned weighted readout. Token states remain SPD-valued until tangent-space classification. Across three EEG benchmarks, SPD-MetaFormer achieves competitive results relative to published Euclidean and manifold baselines. Separate matched reproductions test learned versus uniform weighting within MAtt and GBWAtt. These results suggest that, in the short-sequence and limited-data regimes studied, carefully designed SPD architectures can provide a simpler and effective alternative to adaptive manifold attention.
Problem

Research questions and friction points this paper is trying to address.

brain signal decoding
SPD manifold
attention mechanism
small-data regime
EEG
Innovation

Methods, ideas, or system contributions that make the work stand out.

SPD manifold
attention-free architecture
brain decoding
Fréchet aggregation
small-data regime