Score
Designs and evaluates algorithms, loss functions, and model components that produce and maintain diverse, discriminative feature representations across multiple scales (e.g., feature-pyramid, spatial or temporal scales). This work builds cross-scale variance and covariance constraints, decorrelation or orthogonalization regularizers, and training procedures to prevent feature collapse under shift and to preserve detection- or task-aware discriminability.
This work uncovers the mechanism underlying the power-law decay of prediction error with increasing data volume in deep networks, attributing it to the layerwise recovery of latent compositional features through hierarchical representation learning. Focusing on a class of high-dimensional hierarchical target functions with power-law decaying weights, the authors propose a layerwise spectral algorithm grounded in random matrix theory and resolvent perturbation analysis, establishing—for the first time—a theoretical link between the sequential recovery of features and the global scaling law. By deriving sharp thresholds for individual feature recovery and an explicit power-law form for prediction error, the analysis surpasses conventional gap-dependent perturbation bounds, yielding tight recovery guarantees. Numerical experiments confirm the sequential recovery of features according to their strength, the smoothing of thresholds at finite sample sizes, and superior performance over non-hierarchical kernel methods.
This work investigates the optimality of pretrained feature representations for downstream few-shot learning tasks in transfer learning. Methodologically, it establishes a linear feature transfer model and derives an asymptotic bias–variance decomposition of the downstream risk. Theoretically, it is the first to demonstrate that, on average, the optimal pretrained representation intrinsically exhibits sparsity, and that a phase transition emerges—from hard-thresholding feature selection to soft weighting—without explicit sparse regularization. The analysis integrates asymptotic statistics, multi-task averaging optimization, and linear transfer modeling. The theory precisely characterizes the underlying mechanisms driving both sparsity and the phase transition. Empirical validation on image and text few-shot benchmarks confirms substantial generalization gains over strong baselines. Collectively, this work provides a novel theoretical lens for understanding the intrinsic effectiveness of pretrained representations.
This study investigates the fundamental mechanisms underlying the "grokking" phenomenon—characterized by high training accuracy coupled with delayed generalization—in feature-learning kernels, with a focus on the role of data symmetry. Employing Recursive Feature Machines (RFMs) and iteratively updating feature matrices via the Average Gradient Outer Product (AGOP), the authors analyze grokking behavior in algebraic tasks. Their central finding is that generalization occurs only when the symmetry of the training data is broken. The RFM achieves generalization by recovering the intrinsic group action governing the data, with the learned feature matrix precisely encoding the structure of this symmetry group. This work provides the first empirical evidence that symmetry breaking is a necessary condition for generalization and elucidates the group-theoretic underpinnings of grokking.
This work investigates the training and generalization behavior of high-dimensional ridge regression and random feature models, focusing on scaling laws and renormalization phenomena in the overparameterized regime. Methodologically, it leverages the S-transform from free probability theory—establishing, for the first time, a direct analytical link between the training–generalization gap and spectral properties of the data covariance, thereby deriving a closed-form expression for the generalization error. It identifies feature variance as the dominant factor governing a novel scaling regime and reveals that anisotropic weight distributions induce nontrivial finite-width correction exponents. To absorb statistical fluctuations, the paper introduces a renormalized ridge parameter and constructs a generalized cross-validation–type estimator. The results unify explanations of neural scaling laws, theoretically predict—and empirically validate—multiple power-law behaviors, including benign overfitting and power-law decay of test error. Collectively, this work provides a rigorous analytic framework for high-dimensional statistical learning.
To address over-parameterization and uncertainty in deep networks caused by limited samples in early-stage Alzheimer’s disease (AD) neuroimaging classification, this paper proposes a covariance-driven multi-scale scale-space representation framework. Methodologically, it innovatively integrates covariance modeling and scale-space theory into the neural network backbone, enabling compact representation of high-dimensional brain images and dual-space disentanglement of features and tasks. The framework incorporates scale-space convolution, covariance-based feature extraction, and multi-scale fusion, while supporting gradient-weighted, individualized localization of AD-affected brain regions. Evaluated on the ADNI dataset, the model achieves significantly improved classification accuracy, accelerated convergence, and a substantial reduction in parameter count. Crucially, it retains strong interpretability—precisely identifying AD-specific atrophic brain regions such as the hippocampus and entorhinal cortex.
This work addresses the challenges in nonlinear dimensionality reduction of simultaneously preserving global and local structures and providing interpretability during embedding. We propose a novel embedding method that integrates spectral decomposition with multiscale cross-entropy optimization. By analyzing the influence of spectral modes on embeddings from a graph frequency-domain perspective, our approach uniquely combines spectral basis representations with multiscale nonlinear dimensionality reduction. This integration enables a balanced preservation of both global and local manifold structures while maintaining continuity. To enhance interpretability, we introduce glyph-augmented scatterplots for visual exploration of the embedding process. Quantitative evaluations and case studies demonstrate the effectiveness of our method in achieving structurally faithful and interpretable low-dimensional representations.
This work addresses the degradation of information capacity in multi-granularity embeddings caused by dimensional redundancy and spectral collapse within nested subspaces. To mitigate these issues, the authors propose a self-distillation framework that optimizes embedding geometry through isotropic subspace alignment and introduces a synergistic regularization mechanism combining Soft Collapse Regularization (SCR) and Spectral Isotropy Regularization (SIR). This approach effectively suppresses subspace redundancy while ensuring uniform distribution of low-dimensional prefixes on the hypersphere. Notably, it achieves highly discriminative, semantically dense, and dimensionally flexible representations within a self-distillation paradigm—outperforming existing baselines under high-compression settings while preserving strong information capacity and discriminability.
This work addresses the challenge that existing vision foundation models struggle to effectively capture cross-scale spatial relationships among multimodal satellite images with varying spatial resolutions. To this end, the authors propose Scale-ALiBi, a novel mechanism that incorporates a linear spatial bias—derived from ground sampling distance—into Transformer attention. This is integrated within a joint representation learning framework combining triplet contrastive learning and reconstruction objectives for optical and synthetic aperture radar (SAR) imagery. The key contributions include the first extension of ALiBi to multiscale remote sensing scenarios, a tailored attention mechanism capable of modeling spatial relationships across image patches at different scales, and the creation of the first aligned multimodal, multiscale satellite image dataset. The proposed method achieves significant performance gains on GEO-Bench, and the dataset has been publicly released.
This work addresses the optimization conflicts and semantic error propagation inherent in visual autoregressive (VAR) models for multiscale representation learning, where shared architectures struggle to simultaneously capture global semantics at coarse scales and fine-grained details at high resolutions. To resolve this, the authors propose a scale-aware Mixture-of-Experts (MoE) architecture with token routing that enables scale-adaptive expert selection, thereby decoupling representations across scales. Additionally, they introduce a residual feature alignment strategy tailored to the VAR paradigm, which integrates external self-supervised features to strengthen early-stage semantic modeling. Evaluated on ImageNet at 256×256 resolution, the method significantly outperforms dense baselines in terms of FID, achieving superior performance with fewer parameters and training epochs, and the performance gap widens as training progresses.
Existing object detectors often learn task-driven features that rely on shortcut correlations, failing to adequately capture the underlying annotation structure, which limits their generalization, interpretability, and robustness under task shifts or sparse supervision. To address this, this work proposes an annotation-guided feature enhancement framework that explicitly integrates geometric annotation priors into feature learning for the first time. By constructing a dense spatial feature grid and injecting it into the backbone network—where it fuses with the feature pyramid—the method steers region proposal and detection heads toward representations better aligned with annotation structure. Evaluated on wildlife and remote sensing datasets, the approach significantly improves object focus, reduces background sensitivity, and demonstrates superior generalization and data efficiency in weakly supervised and unseen-task settings.