Score
Designs and implements linear projection–based discriminative classifiers and projection heads that compute directions (often low-dimensional or single-axis) which maximize between-class separation and minimize within-class variance using closed-form LDA solutions. Builds streaming/online variants that update class means and classifier parameters per sample in constant time (sample-wise closed-form updates) and operate without replay buffers or explicit task-boundary signals.
To address the limited discriminative capability of naïve Bayes stemming from its strong conditional independence (isotropic) assumption, this paper proposes Projection Naïve Bayes (PNB), which learns an optimal linear subspace via discriminative projection optimization and performs naïve Bayes factorization of class-conditional densities within this low-dimensional projected space. PNB is the first framework to deeply integrate discriminative projection learning with naïve Bayes modeling, simultaneously enabling dimensionality reduction, visualization, and theoretical interpretability; it is further shown to be equivalent to class-conditional independent component analysis. Extensive experiments across 162 public benchmark datasets demonstrate that PNB significantly outperforms classical probabilistic discriminative models—including Linear Discriminant Analysis (LDA) and Quadratic Discriminant Analysis (QDA)—and matches the accuracy of Support Vector Machines (SVM), while retaining the statistical interpretability and computational efficiency inherent to generative models.
In large-scale few-shot classification, high dimensionality, numerous classes, and extremely limited per-class samples lead to inaccurate covariance estimation and degraded performance of Linear Discriminant Analysis (LDA) and Quadratic Discriminant Analysis (QDA). To address this, we propose a parameter pooling method based on clustering sample covariance matrices. Our approach abandons LDA’s homoscedasticity assumption and employs model-driven spectral regularization for clustering, enabling unified modeling of both singular and non-singular covariances while establishing provable statistical estimation properties. Extensive experiments on synthetic and real-world datasets demonstrate substantial improvements in classification accuracy—particularly under challenging regimes with many classes, very low sample sizes per class, and high dimensionality—outperforming LDA, QDA, and other baselines in robustness and overall performance. The key innovation lies in the first integration of covariance matrix clustering with parameter pooling, jointly ensuring discriminative power, numerical stability, and theoretical interpretability.
This work proposes a geometrically constrained formulation of Deep Linear Discriminant Analysis (Deep LDA) to address performance degradation caused by class cluster overlap or collapse during end-to-end maximum likelihood training. By fixing the class means in the latent space to the vertices of a regular simplex and assuming a shared spherical covariance, the method eliminates degenerate solutions while preserving model simplicity and interpretability. This design enables stable maximum likelihood optimization and yields well-separated class representations. Experimental results on Fashion-MNIST, CIFAR-10, and CIFAR-100 demonstrate that the approach achieves classification accuracy comparable to Softmax baselines, while its latent embeddings exhibit highly structured geometric arrangements in two-dimensional projections.
This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.
Traditional Linear Discriminant Analysis (LDA) is sensitive to noise and fails when the within-class scatter matrix is singular; its stepwise feature selection relies on Wilks’ Λ, which tends to terminate prematurely and degrades discriminative performance. This paper proposes a novel forward discriminant analysis framework. Methodologically, it integrates Pillai’s trace criterion with Uncorrelated LDA (ULDA) for the first time, establishing a unified and interpretable forward feature selection mechanism that avoids premature termination inherent to Wilks’ Λ and naturally accommodates perfectly separable classes. Furthermore, Type I error calibration is incorporated to ensure statistical significance control. Empirical evaluation on both synthetic and real-world datasets demonstrates substantial improvements in classification accuracy and robust false positive rate control, particularly excelling in scenarios of complete class separability.
This work addresses the degeneracy issues in deep Linear Discriminant Analysis (LDA) under maximum likelihood training, which often leads to collapsed class means and covariances, thereby degrading discriminative performance. While cross-entropy training achieves high accuracy, it compromises the probabilistic coherence of the generative model. To reconcile these concerns, the paper introduces a Discriminative Negative Log-Likelihood (DNLL) loss that preserves LDA’s generative structure while incorporating a lightweight penalty on the mixture density to suppress overlap among high-probability regions across classes. This approach enables clear separation in the latent space and represents the first effective integration of generative modeling with discriminative training. Empirical results demonstrate that DNLL attains classification accuracy comparable to Softmax on both synthetic and standard image benchmarks, while significantly improving prediction calibration and feature discriminability.
研究解决了小样本判别分析中因平衡k-shot采样导致的退化问题,通过修正KLPCDA变体并评估其在文本分类任务中的表现。
This work proposes RRLDA-RK, a fast, parameter-free iterative algorithm for reduced-rank linear discriminant analysis (RRLDA) that operates effectively in both classical and high-dimensional settings without relying on strong assumptions or explicit regularization tuning. By integrating techniques from high-dimensional statistics and numerical linear algebra, the method inherently possesses implicit regularization properties and automatically converges to the minimum-norm solution. This ensures theoretical rigor while substantially improving computational efficiency. Empirical evaluations on real high-dimensional datasets demonstrate that RRLDA-RK achieves excellent classification performance alongside strong stability and scalability, addressing the high computational cost typically associated with traditional RRLDA approaches in large-scale, high-dimensional scenarios.
本文针对小样本学习问题,通过系统研究KLPCDA框架在不同场景下的表现,揭示了其核心目标间的交互机制,并提供了选择合适变体的指导原则。
This study addresses the limitations of the original Projection Pursuit Tree (PPtree) classifier, which suffers from shallow tree depth—restricted to fewer splits than the number of classes—and consequently underperforms in high-dimensional, multi-class settings with heterogeneous between-class covariance structures or nonlinear separability. To overcome this, the authors relax the depth constraint and introduce a more flexible class grouping and projection-based splitting mechanism, thereby enhancing the model’s capacity to capture complex decision boundaries. Two novel high-dimensional visualization tools are innovatively designed for diagnostic purposes, and an accompanying R package, PPtreeExt, is developed, featuring an interactive web application that enables side-by-side performance comparison between the original and enhanced classifiers. Empirical evaluations on multiple benchmark high-dimensional datasets demonstrate significant improvements in both classification accuracy and model interpretability.