Score
Design and implement procedures to compute and analyze the empirical spectral density of matrices or operators and to fit relevant distributional parameters (e.g., heavy-tailed shapes). Produce compact, scale‑invariant spectral signatures or fingerprints by aggregating per-component (e.g., per-layer) spectral descriptors for comparison, characterization, or downstream analysis.
This study systematically reviews hyperspectral unmixing methods for accurately estimating land-cover material types, abundances, and spatial distributions from remote sensing imagery. It comprehensively evaluates state-of-the-art algorithms—including sparse representation, low-rank regularization, and deep learning—under both linear and nonlinear mixing models. The suitability of 12 widely used public benchmark datasets is analyzed, and methods are comparatively assessed in terms of unmixing accuracy, robustness to noise and model mismatch, and computational efficiency. Based on this analysis, we propose a scenario-aware method selection guideline. Key limitations are identified: inadequate modeling of endmember variability, poor generalization under limited training samples, and insufficient physical interpretability. Future research directions include integrating physical priors with deep learning architectures and developing generalizable, physics-informed unmixing frameworks. This work provides both theoretical foundations and practical guidance for hyperspectral land-cover composition mapping.
This work addresses the bias in heavy-tailedness estimation arising from aspect-ratio disparities in weight matrices during spectral analysis of deep neural networks. We propose Fixed-Aspect-Ratio Matrix Sampling (FARMS), a method that mitigates this bias by randomly sampling submatrices with fixed aspect ratios, modeling their empirical spectral density (ESD), and fitting α-stable distributions to estimate the tail index. FARMS is the first framework to systematically eliminate the intrinsic aspect-ratio-induced bias in spectral statistics. It exhibits strong cross-architecture and cross-task robustness, significantly improving model diagnostics and layer-wise hyperparameter allocation. Extensive validation across computer vision (CV), scientific machine learning (SciML), and large language model (LLM) pruning tasks confirms its effectiveness: when applied to LLaMA-7B pruning, FARMS reduces perplexity by 17.3%, outperforming state-of-the-art methods.
This paper addresses the challenge of estimating the spectral distribution of high-dimensional weighted sample covariance matrices. We propose WeSpeR, a novel algorithm that (i) rigorously establishes, for the first time, that the limiting spectral distribution admits a regular continuous density; (ii) introduces the first end-to-end joint framework simultaneously performing continuous spectral density estimation, precise support set bounding, and population-level spectral inversion; and (iii) integrates asymptotic random matrix theory analysis, numerical Stieltjes transform inversion, adaptive support localization, and grid optimization. Experiments demonstrate that WeSpeR significantly improves spectral density estimation accuracy, faithfully recovers the true population eigenvalue distribution, and exhibits both theoretical convergence guarantees and strong robustness against model misspecification and noise. By unifying theoretical analysis and practical estimation, WeSpeR establishes a new paradigm for high-dimensional weighted covariance modeling.
Traditional neural networks suffer from limited interpretability and weak theoretical foundations. Method: This paper proposes a novel machine learning paradigm grounded in infinite-dimensional Hilbert spaces, centering on linear operators. It integrates reproducing kernel Hilbert spaces (RKHS), spectral operator learning, wavelet representations, scattering transforms, and Koopman operator theory to formulate learning tasks as sampling, approximation, and dynamical inference in infinite-dimensional function spaces. Contribution/Results: We establish the first unified Hilbert-space-theoretic framework bridging spectral learning and symbolic reasoning. The approach significantly enhances mathematical rigor and model interpretability by grounding learning in well-defined functional-analytic principles. Moreover, it provides a rigorous mathematical foundation and new methodological pathways for deep interdisciplinary integration between signal processing and machine learning—enabling principled analysis of structured data, hierarchical feature extraction, and nonlinear dynamical system modeling.
本文针对大规模稀疏矩阵的频谱分析难题,提出了一种基于GPU的二进制稀疏快速傅里叶变换方法及两种压缩技术,有效减少了内存使用和计算时间。
This work addresses the limitations of conventional supervised dimensionality reduction methods when applied to high-cardinality, sparse, noisy, and class-imbalanced mixed categorical and discrete data. It introduces, for the first time, the density matrix formalism from quantum information theory into this domain, constructing a normalized Gram-type operator based on class-conditional frequencies that satisfies the axioms of a density matrix. This operator yields low-dimensional spectral embeddings in the one-hot encoded space, which are then combined with class-conditional kernel density estimation and a maximum likelihood decision rule for classification. Theoretical analysis reveals a structural property wherein the embedding rank is bounded by the number of classes, and establishes the method’s invariance and stability. Experiments demonstrate that the proposed approach significantly outperforms existing methods on both synthetic and real-world benchmarks, achieving superior robustness and classification performance.
This work addresses the challenges in nonlinear dimensionality reduction of simultaneously preserving global and local structures and providing interpretability during embedding. We propose a novel embedding method that integrates spectral decomposition with multiscale cross-entropy optimization. By analyzing the influence of spectral modes on embeddings from a graph frequency-domain perspective, our approach uniquely combines spectral basis representations with multiscale nonlinear dimensionality reduction. This integration enables a balanced preservation of both global and local manifold structures while maintaining continuity. To enhance interpretability, we introduce glyph-augmented scatterplots for visual exploration of the embedding process. Quantitative evaluations and case studies demonstrate the effectiveness of our method in achieving structurally faithful and interpretable low-dimensional representations.
This work addresses the lack of unified, scalable quantitative methods for lineage tracing, license management, and performance evaluation of large language models (LLMs) by proposing the first model-level signature based on spectral shape. Leveraging heavy-tailed self-regularization theory, the method extracts a compact, data-agnostic, and scale-invariant spectral signature from the empirical spectral density of model weights, exhibiting robustness in post-training settings. Through empirical spectral density analysis, heavy-tailed distribution modeling, and spectral shape metrics, experiments on a large-scale open-source LLM corpus demonstrate that the proposed signature effectively enables unsupervised clustering, lineage attribution, and performance trend prediction, thereby validating its practical utility and effectiveness for managing large models.
Traditional clustering methods struggle with high-dimensional data where different subsets of features correspond to distinct cluster structures and some features are uninformative. This work proposes a local spectral clustering framework that formulates local clustering as a “clustering of clusterings” problem. It groups features via a label-invariant clustering matrix and constructs feature-specific Gaussian kernel similarity matrices based on a heterogeneous sub-Gaussian mixture model. The approach jointly identifies homogeneous feature groups and their corresponding sample partitions without requiring explicit likelihood evaluation or Bayesian inference. Experiments demonstrate that the method effectively uncovers complex heterogeneous clustering structures in both synthetic and real-world datasets, exhibiting superior performance and practical utility.
本文研究了基于顺序周期图的谱密度积分估计量的自归一化方法,解决了线性和非线性泛函中未知谱量的问题。
This work addresses the gap between algorithmic prototypes and efficient implementations in scientific research by proposing a lightweight approach to translate statistical and machine learning algorithms—such as kernel ridge regression and stochastic gradient descent matrix factorization—from mathematical formulations into readable, high-performance C++ code. Leveraging the Eigen template library for core linear algebra operations—including kernel matrix construction, regularized solvers, and vectorized updates—the implementation seamlessly integrates into the Python ecosystem via pybind11, enabling efficient interoperability with NumPy arrays. The project provides concise, reproducible code examples that encapsulate common computational patterns in research, significantly lowering the barrier for researchers to adopt C++ for high-performance development while balancing performance, readability, and usability.