Score
Designs and evaluates linear projection bases or directions that use empirical Fisher (gradient-statistics) information to emphasize target- or class-specific signal and suppress shared or background structure. This involves estimating empirical Fisher matrices or gradient-statistics, constructing Fisher-weighted or logit-weighted projections, and solving (regularized) generalized eigenproblems to produce discriminative bases or projection operators.
Classification of high-dimensional matrix-valued data—common in neuroimaging (e.g., fMRI, EEG) and signal processing—remains challenging due to structural dependencies and limited sample sizes. Method: This paper proposes a nonparametric linear discriminant analysis (NPLDA) specifically designed for matrix data. It introduces a nonparametric empirical Bayes framework coupled with nonparametric maximum likelihood estimation (NPMLE) into matrix LDA, eliminating reliance on parametric assumptions about prior distributions. Leveraging the matrix normal distribution, NPLDA preserves structural information via vectorization and standardization, while jointly estimating nonparametric distributions of class-conditional means and covariance matrices. Contribution/Results: Extensive experiments on multiple real neuroimaging datasets and synthetic benchmarks demonstrate that NPLDA consistently outperforms conventional vectorized LDA, parametric matrix LDA, and deep learning baselines, achieving average accuracy gains of 3.2–7.8%. Moreover, it exhibits superior robustness and adaptability in small-sample and high-dimensional sparse regimes.
To address the limited discriminative capability of naïve Bayes stemming from its strong conditional independence (isotropic) assumption, this paper proposes Projection Naïve Bayes (PNB), which learns an optimal linear subspace via discriminative projection optimization and performs naïve Bayes factorization of class-conditional densities within this low-dimensional projected space. PNB is the first framework to deeply integrate discriminative projection learning with naïve Bayes modeling, simultaneously enabling dimensionality reduction, visualization, and theoretical interpretability; it is further shown to be equivalent to class-conditional independent component analysis. Extensive experiments across 162 public benchmark datasets demonstrate that PNB significantly outperforms classical probabilistic discriminative models—including Linear Discriminant Analysis (LDA) and Quadratic Discriminant Analysis (QDA)—and matches the accuracy of Support Vector Machines (SVM), while retaining the statistical interpretability and computational efficiency inherent to generative models.
How can the diagonal of the Fisher Information Matrix (FIM) be approximated at zero computational cost? This paper introduces Squisher, the first method to systematically demonstrate that the squared-gradient moving averages—already maintained by adaptive optimizers such as Adam—serve as high-fidelity, zero-cost approximations to the FIM diagonal, requiring no additional forward/backward passes or stochastic sampling. By reusing existing training-time statistics, Squisher enables parameter sensitivity quantification without increasing computational overhead. Extensive experiments across five canonical tasks—including network pruning, uncertainty estimation, and continual learning—show that Squisher matches the accuracy of standard FIM diagonal estimation while substantially outperforming existing baselines. Moreover, it is fully plug-and-play, integrating seamlessly into standard training pipelines without architectural or optimization modifications.
Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.
This work addresses the challenge of Bayesian dimensionality reduction in gradient-free settings—such as purely data-driven or simulation-based inference—where gradients of the target posterior are unavailable. To this end, we propose a score-ratio-based Bayesian dimensionality reduction framework that operates without explicit gradient computation. Our method innovatively extends score matching to gradient-deficient scenarios for dimensionality reduction; introduces a dedicated neural network architecture that jointly learns the score-ratio function and identifies an informative low-dimensional subspace; and incorporates manifold-aware regularization alongside an iterative basis-optimization algorithm leveraging eigenvalue truncation. Evaluated on PDE-constrained Bayesian inverse problems and conditional generation tasks, the approach achieves high-fidelity low-dimensional posterior approximations using only limited simulation data, significantly outperforming standard score matching and other gradient-free dimensionality reduction baselines.
本文提出了一种基于最近邻关系增强的数据降维方法,通过构建和分解能够捕捉局部协方差结构的矩阵来寻找多变量数据的有趣投影,该方法有效估计了密度信息矩阵。
This work addresses the lack of theoretical characterization for discriminative subspaces under orthogonality constraints in multi-label linear discriminant analysis, particularly concerning objective equivalence, effective dimensionality, and estimation error. Focusing on multi-label Fisher discriminant analysis with Stiefel orthogonality constraints, we integrate algebraic structure and statistical analysis to reveal the rank properties of the between-class scatter matrix and establish conditions under which four Fisher objectives become equivalent. We further prove that the effective discriminative dimension can exceed the classical \(C-1\) upper bound. Our analysis yields nearly minimax-optimal subspace estimation error bounds—upper bound \(O(k_{\max}\sqrt{d\log d/n}/\text{gap}_r)\) and lower bound \(\Omega(\sigma^2 d/(n\cdot\text{gap}_r))\)—and provides the first distance preservation and robustness guarantees for the multi-label setting. Numerical experiments validate key algebraic identities and the influence of multi-label-specific parameters such as \(k_{\max}\) and \(\kappa(S_t^{\text{ML}})\).
This study addresses the limitation that coefficient sparsity in Kolmogorov–Arnold Networks (KANs) cannot directly serve as a pruning criterion, and investigates the unclear alignment between architectural simplicity and statistical Fisher simplicity. By comparing the Fisher null space properties of dead ReLUs and fixed-basis KANs, this work proposes an effective edge metric based on graph paths, supported by mathematical derivations and controlled experiments integrating Fisher information matrix theory with basis Gram matrix analysis. The findings reveal that Fisher simplicity in multi-layer KANs is governed by data propagation paths rather than coefficient magnitudes. Furthermore, zero effective exposure is shown to accurately identify Fisher null directions, demonstrating that individual coefficient magnitudes are insufficient as a Fisher-based pruning criterion for KANs. These insights refine the theoretical understanding of model simplification.
This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.
This study addresses the evaluation and selection of source-level likelihood ratio (LR) systems for forensic evidence-to-reference comparison tasks by proposing an integrated analytical framework that balances performance and practical feasibility. The authors employ strictly proper scoring rules to quantify how effectively each system updates Bayesian prior odds and present the first systematic comparison among specific-source feature-based, common-source anchored, and unanchored score-based LR approaches. Their findings reveal that specific-source feature-based LRs achieve the highest performance but incur substantial experimental costs, whereas common-source feature-based methods offer strong discriminative power with significantly reduced implementation complexity. All LR systems substantially outperform a baseline relying solely on prior odds. This work thus provides both theoretical grounding and practical guidance for selecting LR systems in forensic practice.