Score
Designs and implements deep transformer-based classifiers that ingest spectral measurements and output categorical labels, including architectures, training procedures, and latent-feature encoders tailored for spectral inputs. Builds and evaluates models and pipelines that learn discriminative latent representations from spectra and assess robustness and generalization (for example across replicates or data shifts).
This paper investigates whether Transformers, under unsupervised pretraining, can autonomously learn spectral algorithms without task-specific supervision. Method: We introduce an algorithm-unfolding modeling paradigm—distinct from in-context learning—that formalizes how Transformers implicitly acquire algorithmic knowledge through experience-like training. Leveraging spectral analysis and Gaussian mixture model (GMM) theory, we provide constructive theoretical proofs that multi-layer Transformers can implement principal component analysis (PCA) and GMM-based clustering, without prompting or downstream fine-tuning. Contribution/Results: We establish, for the first time, a provable correspondence between the Transformer architecture and classical spectral methods. Our analysis yields convergence guarantees, and empirical evaluation on both synthetic and real-world datasets demonstrates performance approaching statistical optimality. The work reveals an intrinsic algorithm-learning mechanism within Transformers, bridging architectural design and principled unsupervised algorithm discovery.
Hyperspectral image (HSI) classification faces critical challenges including severe label scarcity, excessively high spectral dimensionality, prohibitive computational overhead, and poor intrinsic interpretability. Method: This paper systematically reviews over 300 peer-reviewed works published before 2025 and proposes the first end-to-end HSI-Transformer methodology stack, encompassing spatial-spectral tokenization, adaptive positional encoding, lightweight multi-head attention, robust feature extraction, and interpretability-aware loss design. Contribution/Results: We introduce a “property–architecture” alignment analytical framework that explicitly identifies four fundamental technical gaps. Furthermore, we establish the inaugural methodological framework for HSI-Transformer classification, providing both theoretical foundations and practical guidelines for edge deployment, cross-domain generalization, and inherent model interpretability.
This work addresses the lack of mechanistic understanding, principled layer selection, and theoretically grounded feature alignment in vision transformer (ViT) knowledge distillation. We first discover that CaiT and Swin exhibit similar spectral encoding patterns, motivating a novel interpretable knowledge distillation framework based on spectral analysis—SpectralKD. Our method leverages frequency-domain spectral analysis to uncover intrinsic information transfer dynamics between teacher and student models, thereby guiding the selection of critical feature layers and enabling cross-layer spectral alignment. Evaluated on ImageNet-1k, SpectralKD improves Top-1 accuracy by 5.2% for DeiT-Tiny and 1.4% for Swin-Tiny student models. Moreover, distilled students exhibit significantly converged spectral characteristics toward their teachers, demonstrating both substantial performance gains and enhanced interpretability.
This work addresses the parameter redundancy in embedding dimensions of large language models, which incurs substantial computational and memory costs during scaling. The authors propose L-Transformer, the first architecture to enable differentiable spectral decomposition of the embedding space. By leveraging the third-order tensor L-product, token embeddings are reshaped into spectral tensor slices, and attention and feed-forward operations are performed in the transform domain. This yields a Tensor Transformer composed of p independent spectral sub-transformers, introducing an inductive bias at the frequency level that supports slice-dependent frequency scaling to enhance generalization while remaining compatible with standard training pipelines. Experiments show that on IMDB and AG News, the encoder achieves up to 75% parameter reduction (with p=4) while maintaining competitive accuracy, and fully recovers baseline performance at BERT-base width.
Existing random feature methods for linearizing Transformer attention lack systematic kernel function design and evaluation. Method: This paper proposes Spectraformer—the first comparable, scalable, unified framework that explicitly models kernel function structure in linearized attention. By jointly designing nonlinear component functions (e.g., ReLU, exp, sin) and structured or random weight matrices, it systematically uncovers the complementarity and task dependency of different kernels for long-range text modeling. Contribution/Results: Evaluated on all three text-based Long-Range Arena (LRA) benchmarks, Spectraformer demonstrates that kernel selection critically impacts accuracy—achieving significant improvements while preserving linear time and space complexity. The framework is fully open-sourced.
The internal mechanisms of large language models lack systematic interpretability analysis. Method: This paper introduces CAST, the first framework to directly estimate the actual transformation matrices of individual Transformer layers via the Moore–Penrose pseudoinverse, enabling probe-free, fine-grained full-spectrum analysis based on six spectral metrics. Integrating CKA similarity matrices with kernel analysis techniques, CAST characterizes fundamental architectural differences between encoders and decoders. Contribution/Results: We reveal that decoders exhibit a “compression–expansion” cyclic dynamic, whereas encoders preserve high-rank feature representations. Transformer layers are functionally partitioned into three distinct phases: feature extraction, compression, and specialization. CAST establishes a verifiable, generalizable paradigm for analyzing inter-layer information flow and functional division of labor in Transformers, advancing mechanistic interpretability beyond heuristic probing.
This study investigates how linear recurrent Transformers equipped with layer normalization implicitly learn the power method through gradient descent when trained on principal component prediction tasks. The work reveals an “algorithmic implicit bias”: in the absence of explicit supervision, the self-attention layers automatically converge to solutions that implement power iterations, with each layer corresponding to one update step of the power method. Theoretical analysis demonstrates that layer normalization is essential for realizing the exact power method—models without it fail to replicate the algorithm, resulting in significantly degraded performance. This paper is the first to establish the pivotal role of layer normalization in inducing algorithmic inductive biases and provides provable guarantees for the resulting performance gap.
This work addresses the opacity of Transformer inference in separator-free multiclass linear classification by imposing permutation equivariance constraints between features and labels. This constraint enforces a highly structured weight configuration while preserving functional equivalence to the original model. Leveraging this design, the study presents the first explicit inter-layer recursive update rule extracted from an end-to-end trained Softmax-based Transformer, thereby uncovering the implicit geometric algorithmic nature of the attention mechanism. The proposed approach not only enhances class separability but also theoretically guarantees desired class-pair robustness, offering both interpretability and performance benefits in linearly separable classification settings.
This paper presents a novel approach, Spectral-Interpretable and -Enhanced Transformer (SIEFormer), which leverages spectral analysis to reinterpret the attention mechanism within Vision Transformer (ViT) and enhance feature adaptability, with particular emphasis on challenging Generalized Category Discovery (GCD) tasks. The proposed SIEFormer is composed of two main branches, each corresponding to an implicit and explicit spectral perspective of the ViT, enabling joint optimization. The implicit branch realizes the use of different types of graph Laplacians to model the local structure correlations of tokens, along with a novel Band-adaptive Filter (BaF) layer that can flexibly perform both band-pass and band-reject filtering. The explicit branch, on the other hand, introduces a Maneuverable Filtering Layer (MFL) that learns global dependencies among tokens by applying the Fourier transform to the input ``value"features, modulating the transformed signal with a set of learnable parameters in the frequency domain, and then performing an inverse Fourier transform to obtain the enhanced features. Extensive experiments reveal state-of-the-art performance on multiple image recognition datasets, reaffirming the superiority of our approach through ablation studies and visualizations.
This study addresses the issue of erroneous invariance attribution in preprocessing for spectral foundation models by proposing a decoupled evaluation paradigm that distinguishes the contributions of normalization from representation learning. Through Raman spectroscopy analysis, controlled experiments, and numerical validation, we demonstrate that the invariance observed in multimodal models primarily stems from parameter-free preprocessing normalization rather than learned representations. These findings reveal that current models fail to surpass simple normalization baselines, effectively correcting prevailing cognitive biases regarding spectral representation learning capabilities within the field. Consequently, this work establishes a new benchmark for evaluation methodologies, emphasizing the necessity of rigorously isolating preprocessing effects when assessing the true efficacy of spectral foundation models.