fisher discriminative projection

Designs and evaluates linear projection bases or directions that use empirical Fisher (gradient-statistics) information to emphasize target- or class-specific signal and suppress shared or background structure. This involves estimating empirical Fisher matrices or gradient-statistics, constructing Fisher-weighted or logit-weighted projections, and solving (regularized) generalized eigenproblems to produce discriminative bases or projection operators.

fisherdiscriminativeprojection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Nonparametric Linear Discriminant Analysis for High Dimensional Matrix-Valued Data

Jul 25, 2025
SO
Seungyeon Oh
🏛️ Sookmyung Women’s University | Sungshin Women’s University | Data Science Center at Sungshin Women’s University | Research Institute of Natural Science, Sookmyung Women’s University

Classification of high-dimensional matrix-valued data—common in neuroimaging (e.g., fMRI, EEG) and signal processing—remains challenging due to structural dependencies and limited sample sizes. Method: This paper proposes a nonparametric linear discriminant analysis (NPLDA) specifically designed for matrix data. It introduces a nonparametric empirical Bayes framework coupled with nonparametric maximum likelihood estimation (NPMLE) into matrix LDA, eliminating reliance on parametric assumptions about prior distributions. Leveraging the matrix normal distribution, NPLDA preserves structural information via vectorization and standardization, while jointly estimating nonparametric distributions of class-conditional means and covariance matrices. Contribution/Results: Extensive experiments on multiple real neuroimaging datasets and synthetic benchmarks demonstrate that NPLDA consistently outperforms conventional vectorized LDA, parametric matrix LDA, and deep learning baselines, achieving average accuracy gains of 3.2–7.8%. Moreover, it exhibits superior robustness and adaptability in small-sample and high-dimensional sparse regimes.

Classify matrix-valued data in neuroimaging and signal processingExtend Fisher's LDA for matrix observations using matrix normal distributionImprove classification via nonparametric empirical Bayes and NPMLE

Optimal Projections for Classification with Naive Bayes

Sep 09, 2024
DP
David P. Hofmeyr
🏛️ Lancaster University | Swiss Data Science Center | EPFL | Kohort

To address the limited discriminative capability of naïve Bayes stemming from its strong conditional independence (isotropic) assumption, this paper proposes Projection Naïve Bayes (PNB), which learns an optimal linear subspace via discriminative projection optimization and performs naïve Bayes factorization of class-conditional densities within this low-dimensional projected space. PNB is the first framework to deeply integrate discriminative projection learning with naïve Bayes modeling, simultaneously enabling dimensionality reduction, visualization, and theoretical interpretability; it is further shown to be equivalent to class-conditional independent component analysis. Extensive experiments across 162 public benchmark datasets demonstrate that PNB significantly outperforms classical probabilistic discriminative models—including Linear Discriminant Analysis (LDA) and Quadratic Discriminant Analysis (QDA)—and matches the accuracy of Support Vector Machines (SVM), while retaining the statistical interpretability and computational efficiency inherent to generative models.

Enhancing discriminatory power through alternative basis factorisationFinding optimal linear projections for Naive Bayes classificationPerforming projection pursuit with multinomial likelihood optimization

How can the diagonal of the Fisher Information Matrix (FIM) be approximated at zero computational cost? This paper introduces Squisher, the first method to systematically demonstrate that the squared-gradient moving averages—already maintained by adaptive optimizers such as Adam—serve as high-fidelity, zero-cost approximations to the FIM diagonal, requiring no additional forward/backward passes or stochastic sampling. By reusing existing training-time statistics, Squisher enables parameter sensitivity quantification without increasing computational overhead. Extensive experiments across five canonical tasks—including network pruning, uncertainty estimation, and continual learning—show that Squisher matches the accuracy of standard FIM diagonal estimation while substantially outperforming existing baselines. Moreover, it is fully plug-and-play, integrating seamlessly into standard training pipelines without architectural or optimization modifications.

Comparing Squisher performance with Fisher diagonal and baselinesEstimating Fisher diagonal without extra computational costReusing squared gradient accumulator for Fisher approximation

Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing

Aug 28, 2023
YZ
Yihan Zhang
🏛️ Institute of Science and Technology Austria | University of Cambridge

Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.

Characterizing spectral estimators for correlated Gaussian designsEstimating parameters in high-dimensional generalized linear modelsIdentifying optimal preprocessing for efficient parameter estimation

Dimension reduction via score ratio matching

Oct 25, 2024
RB
R. Baptista
🏛️ California Institute of Technology | Massachusetts Institute of Technology

This work addresses the challenge of Bayesian dimensionality reduction in gradient-free settings—such as purely data-driven or simulation-based inference—where gradients of the target posterior are unavailable. To this end, we propose a score-ratio-based Bayesian dimensionality reduction framework that operates without explicit gradient computation. Our method innovatively extends score matching to gradient-deficient scenarios for dimensionality reduction; introduces a dedicated neural network architecture that jointly learns the score-ratio function and identifies an informative low-dimensional subspace; and incorporates manifold-aware regularization alongside an iterative basis-optimization algorithm leveraging eigenvalue truncation. Evaluated on PDE-constrained Bayesian inverse problems and conditional generation tasks, the approach achieves high-fidelity low-dimensional posterior approximations using only limited simulation data, significantly outperforming standard score matching and other gradient-free dimensionality reduction baselines.

Extends gradient-based dimension reduction to gradient-unavailable problemsImproves accuracy of low-dimensional basis identification with limited dataLearns score ratio function for diagnostic matrices without gradients

Latest Papers

What's happening recently
View more

This work addresses the lack of theoretical characterization for discriminative subspaces under orthogonality constraints in multi-label linear discriminant analysis, particularly concerning objective equivalence, effective dimensionality, and estimation error. Focusing on multi-label Fisher discriminant analysis with Stiefel orthogonality constraints, we integrate algebraic structure and statistical analysis to reveal the rank properties of the between-class scatter matrix and establish conditions under which four Fisher objectives become equivalent. We further prove that the effective discriminative dimension can exceed the classical \(C-1\) upper bound. Our analysis yields nearly minimax-optimal subspace estimation error bounds—upper bound \(O(k_{\max}\sqrt{d\log d/n}/\text{gap}_r)\) and lower bound \(\Omega(\sigma^2 d/(n\cdot\text{gap}_r))\)—and provides the first distance preservation and robustness guarantees for the multi-label setting. Numerical experiments validate key algebraic identities and the influence of multi-label-specific parameters such as \(k_{\max}\) and \(\kappa(S_t^{\text{ML}})\).

Multilabel Fisher DiscriminantOrthogonal ConstraintsScatter Matrix

This study addresses the limitation that coefficient sparsity in Kolmogorov–Arnold Networks (KANs) cannot directly serve as a pruning criterion, and investigates the unclear alignment between architectural simplicity and statistical Fisher simplicity. By comparing the Fisher null space properties of dead ReLUs and fixed-basis KANs, this work proposes an effective edge metric based on graph paths, supported by mathematical derivations and controlled experiments integrating Fisher information matrix theory with basis Gram matrix analysis. The findings reveal that Fisher simplicity in multi-layer KANs is governed by data propagation paths rather than coefficient magnitudes. Furthermore, zero effective exposure is shown to accurately identify Fisher null directions, demonstrating that individual coefficient magnitudes are insufficient as a Fisher-based pruning criterion for KANs. These insights refine the theoretical understanding of model simplification.

Fisher simplicityinterpretabilityKolmogorov-Arnold Networks

This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.

calibrated inferenceeigenspace perturbationerror propagation

This study addresses the evaluation and selection of source-level likelihood ratio (LR) systems for forensic evidence-to-reference comparison tasks by proposing an integrated analytical framework that balances performance and practical feasibility. The authors employ strictly proper scoring rules to quantify how effectively each system updates Bayesian prior odds and present the first systematic comparison among specific-source feature-based, common-source anchored, and unanchored score-based LR approaches. Their findings reveal that specific-source feature-based LRs achieve the highest performance but incur substantial experimental costs, whereas common-source feature-based methods offer strong discriminative power with significantly reduced implementation complexity. All LR systems substantially outperform a baseline relying solely on prior odds. This work thus provides both theoretical grounding and practical guidance for selecting LR systems in forensic practice.

likelihood-ratio systemsperformance vs. feasibilityscoring rules

Hot Scholars

YY

Yan Yan

University of Illinois Chicago
Computer VisionMultimediaMachine Learning
ZL

Zhiling Lan

Professor of Computer Science, University of Illinois Chicago
cluster schedulingenergy efficiencyAI4Sysmodeling and simulation
SS

Shun Shao

University of Cambridge
Artificial IntelligenceMachine learning
CR

Chris Russell

Associate Professor, University of Oxford
Ethical Machine LearningComputer VisionOptimisationEthical AI
GL

Gaowen Liu

Cisco Research
machine learningcomputer visionmultimedia.