Score
Designs and evaluates algorithms that transform spectral inputs into low-dimensional embeddings or feature representations that suppress or remove batch-specific variation, using spectral-embedding techniques together with explicit batch-effect correction to produce domain-invariant features. Builds and analyzes training objectives and modules—for example adversarial or entropy-regularized invariance losses, alignment or projection steps—to improve cross-batch generalization and robustness of downstream models.
Traditional spectral embedding (SE) methods suffer from three key limitations: poor generalizability to out-of-sample nodes, computational intractability on large-scale graphs, and insufficient separability of learned eigenvectors—hindering downstream clustering and visualization. To address these, we propose GrEASE, the first end-to-end differentiable and generalizable deep spectral embedding framework. Its core contributions are: (1) a differentiable neural architecture approximating the graph Laplacian operator, enabling generalizable SE learning; (2) a feature-decoupling regularization mechanism that explicitly enhances geometric separability of embeddings; and (3) NUMAP—the first generalizable variant of UMAP. Experiments demonstrate that GrEASE consistently approximates ground-truth spectral embeddings across diverse datasets, enables real-time out-of-sample embedding, and significantly improves both the generalizability and efficiency of UMAP. The implementation is publicly available.
This work addresses training instability and inefficiency in contrastive learning caused by excessively large gradient norms. We propose a spectral-domain analysis framework for gradient stability, establishing the first non-asymptotic spectral band constraint and proving an $O(1/ au^2)$ upper bound on the gradient norm—revealing the joint influence of batch spectral diversity, feature alignment, and temperature $ au$. Based on this theory, we design Greedy-64, a spectral-aware greedy batch selection algorithm that uses effective rank to quantify feature anisotropy for efficient batch construction. We further integrate batch whitening to suppress gradient variance. Experiments show that Greedy-64 accelerates training by 15% on ImageNet-100 while maintaining consistent accuracy gains on CIFAR-10; batch whitening reduces gradient variance by 1.37×, empirically validating the theoretical bound.
This work addresses the challenge of quantifying input feature importance in deep neural networks. We propose a training-embedded spectral reparameterization method that directly employs the eigenvalues associated with input nodes as robust proxies for feature relevance, enabling simultaneous feature importance estimation and model training—without post-hoc analysis or auxiliary supervision. Our key contribution is the first use of input-node eigenvalue sensitivity in spectral neural networks to characterize relative feature importance, coupled with spectral reparameterization during optimization to ensure numerical stability. Experiments on both synthetic and real-world datasets demonstrate that the method significantly improves feature selection efficiency and model interpretability while strictly preserving predictive accuracy—achieving zero performance degradation.
This paper investigates the generalization performance of random feature methods under generalized spectral regularization—including explicit schemes (e.g., Tikhonov regularization) and implicit schemes (e.g., gradient descent and accelerated algorithms). Leveraging insights from neural tangent kernel (NTK) theory, it establishes optimal learning rates for the first time under non-RKHS-type regularity assumptions, thereby unifying the analysis of implicit and explicit regularization mechanisms. Methodologically, the work integrates random feature mappings, spectral analysis, source condition modeling, and rigorous generalization error bound derivation to obtain tight convergence rates for both neural networks and neural operators. Key contributions are: (1) extending optimal convergence guarantees beyond the RKHS framework to broader classes of spectral regularity; (2) providing a unified, tight, and improved theoretical bound applicable to diverse kernel-based algorithms; and (3) substantially enhancing computational efficiency and depth of generalization analysis for large-scale kernel methods.
This work addresses the limited generalization of hyperspectral image models caused by spectral configuration discrepancies across sensors—such as variations in wavelength coverage, band sampling, and channel dimensions—and proposes LESSViT, a sensor-flexible Vision Transformer architecture. LESSViT enables efficient explicit spatial-spectral joint modeling via low-rank decomposition, supports arbitrary spectral inputs through channel-agnostic patch embedding and wavelength-aware positional encoding, and significantly reduces computational complexity with a novel LESS Attention mechanism. Furthermore, the authors introduce HyperMAE, a pretraining strategy that leverages decoupled spatial-spectral masking and hierarchical channel sampling. Evaluated on the SpectralEarth benchmark, LESSViT maintains strong in-domain performance while substantially improving robustness to spectral shifts, demonstrating its effectiveness for scalable and generalizable hyperspectral representation learning.
This study investigates the vulnerability mechanisms of vision-language models under adversarial attacks, with a focus on how the spectral structure of intermediate linear transformations influences model robustness. To this end, the authors propose a white-box Spectral Subspace-Guided Attack (SSGRA), which enhances attack efficacy by aligning intermediate representations with the subspace spanned by right singular vectors. This work is the first to reveal the adversarial fragility of vision-language models from the perspective of spectral subspaces, achieving higher attack success rates than existing baselines. Moreover, it offers novel theoretical insights and a principled technical pathway toward understanding and improving the robustness of such models.
This work addresses the challenge of directly transferring pretrained RGB vision models to hyperspectral image analysis, where a fundamental mismatch exists between the three-channel input assumption and the high-dimensional spectral nature of hyperspectral data. To overcome this, the authors propose a novel partially trainable tensor decomposition strategy that decouples pretrained convolutional kernels into spatial and spectral components. The original three-channel spectral part is replaced with a high-dimensional, learnable spectral component, thereby constructing new filters tailored for hyperspectral inputs. This approach uniquely integrates trainable tensor decomposition into transfer learning, preserving the powerful spatial feature extraction capabilities of the original model while effectively modeling hyperspectral characteristics. Extensive experiments demonstrate that the proposed method significantly outperforms existing transfer learning approaches across multiple hyperspectral datasets, achieving both higher accuracy and improved robustness.