train classification models

Designs, trains, and fine‑tunes classification systems and related learning components—classifiers, diagnostic and auxiliary classifiers, discriminators, discriminant functions, decision trees, kernel learners, and procedures for incremental or inductive updates—covering small-data fine‑tuning through large‑pretrained and large‑model training. Builds and integrates classifier guidance or classifier‑free guidance with differentiable generators, implements training workflows and monitoring (training logs, anomaly detection), and analyzes model behavior via generalization evaluation, failure‑case identification, and synthetic tests for leakage.

trainclassificationmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$220K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Efficient Large-Scale Learning of Minimax Risk Classifiers

Nov 18, 2025
KB
Kartheek Bondugula
🏛️ Basque Center for Applied Mathematics (BCAM) | IKERBASQUE, Basque Foundation for Science

Training minimax risk classifiers (MRCs) for large-scale multi-class problems is computationally prohibitive; existing stochastic subgradient methods fail to effectively optimize the max-expected-loss objective. Method: We propose a deterministic optimization framework integrating constraint generation and column generation, eliminating stochastic approximations. It iteratively expands both the constraint set (samples) and the category subset (classes), enabling scalable optimization over large datasets and high-dimensional label spaces. Contribution/Results: This is the first deterministic algorithm supporting large-scale multi-class MRC training. On multiple benchmark datasets, it achieves 10–100× speedup over conventional methods—with acceleration increasing as the number of classes grows—while strictly preserving the MRC’s robustness guarantees and classification accuracy.

Accelerating classification with multiple classes using constraint generationEfficient large-scale learning for minimax risk classifiersOvercoming limitations of stochastic subgradient methods

Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective

Mar 13, 2025
XZ
Xiaoming Zhao
🏛️ University of Illinois Urbana-Champaign

The mechanistic role of classifier-free guidance (CFG) in conditional generation remains poorly understood, hindering principled improvements. Method: This work introduces a classifier-centric unified analytical framework, revealing that both classifier guidance and CFG fundamentally steer denoising trajectories away from classification decision boundaries to enforce conditional control. Leveraging this insight, the authors propose the first flow-matching-based general post-processing method to calibrate pre-trained diffusion models’ generation bias near decision boundaries—without model fine-tuning or training modifications. Contribution/Results: Extensive experiments across multiple image datasets demonstrate that the method significantly enhances class-discriminative robustness and detail fidelity of generated samples. It establishes a novel paradigm for interpreting and improving CFG, offering both theoretical insight and a practical, plug-and-play tool for enhancing conditional generation quality.

Exploring the role of classifiers in conditional generation.Improving distribution alignment near decision boundaries.Understanding classifier-free guidance in diffusion models.

Learning accurate and interpretable tree-based models

May 24, 2024
MB
Maria-Florina Balcan
🏛️ Carnegie Mellon University

In same-domain repeated access scenarios, decision trees and their ensembles (e.g., random forests, GBDTs) struggle to jointly optimize accuracy and interpretability. Method: This paper proposes a learnable, tunable unified tree framework. It introduces (1) a parameterized splitting criterion that continuously interpolates between entropy and Gini impurity, enabling data-adaptive optimal splits; (2) a theoretical characterization of sample complexity, providing generalization guarantees for the interpretability–accuracy trade-off; and (3) joint optimization of Bayesian decision trees, minimum-cost-complexity pruning, and ensemble hyperparameters. Results: Extensive experiments on real-world datasets demonstrate that the framework significantly improves the consistency between predictive accuracy and model interpretability. It offers both theoretical rigor—via provable generalization bounds—and practical utility—through end-to-end differentiability and seamless integration into existing tree-based pipelines. The approach bridges a critical gap between statistical performance and human-understandable structure in tree learning.

Develop data-specific tree-based learning algorithmsOptimize explainability versus accuracy trade-offTune hyperparameters in pruning and ensembles

This work addresses fundamental challenges in supervised learning—namely, the theoretical limitations of function approximation, the difficulty of functional enhancement across domains in transfer learning, and the trade-off between efficiency and accuracy in active learning. To tackle these issues, we propose a novel unified framework that integrates manifold learning, function boosting theory, and signal separation techniques, operating without explicit training mechanisms. The resulting approach enables efficient function approximation, rigorous transferability analysis, and rapid classification. Empirical evaluations demonstrate that the proposed algorithm achieves state-of-the-art classification accuracy while substantially improving computational efficiency, thereby offering robust theoretical foundations for both supervised and transfer learning paradigms.

active learningclassificationfunction approximation

This paper addresses the fundamental trade-off between zero false negatives and low false positives in dynamic classification. To resolve this, we propose a lightweight multi-model collaborative framework. Methodologically, we introduce a novel self-supervised classification learning mechanism; dynamically partition input data into $N$ mutually exclusive subsets; train independent submodels for parallel prediction; and incorporate a confidence-threshold-based filtering and prediction rejection mechanism—eliminating unreliable predictions without requiring auxiliary verification models. Supervised feedback is further leveraged to iteratively refine model performance. Experiments demonstrate strict zero false negatives and a 37.2% reduction in false positive rate over state-of-the-art ensemble methods under low partitioning error; under high partitioning error, the framework maintains robustness comparable to current best models. Our core contribution is a reliability- and efficiency-aware lightweight paradigm for dynamic classification.

Achieve zero missed detections and minimal false positivesEnhance prediction accuracy via dynamic classificationImprove model efficiency with self-supervised learning

Latest Papers

What's happening recently
View more

This study addresses classification problems commonly characterized by missing data, the need to incorporate expert prior knowledge, and demands for interpretable decisions. The authors propose an expert-guided class-conditional modeling approach that constructs interpretable goodness-of-fit features to quantify the consistency between incomplete observations and expert-derived models. By integrating these features with a small set of transparent summary statistics, they design a lightweight yet effective discriminative classifier. This method embeds domain knowledge directly into the class-conditional generative process, achieving substantially improved classification performance under limited sample sizes. Evaluated on seismic monitoring tasks, the system functions as a transparent screening tool that effectively reduces expert workload and outperforms mainstream machine learning methods.

expert knowledgegoodness-of-fitinformative missingness

This study addresses the challenge of detecting anomalous events in large-scale, high-voltage power grid operational data by systematically evaluating the performance of neural networks, k-nearest neighbors, support vector machines, and unsupervised learning methods under complex contextual conditions. The findings reveal that grid anomalies exhibit strong context dependency. Among the evaluated approaches, neural networks significantly outperform traditional methods in overall detection accuracy, while unsupervised learning algorithms demonstrate superior robustness and efficiency in scenarios involving concurrent multiple anomalies. This work not only validates the advantages of deep learning for anomaly detection in power systems but also highlights the practical value of unsupervised methods when labeled data are scarce and fault patterns are intricately coupled.

Anomaly DetectionMachine LearningOperational Data

Industrial anomaly detection is often hindered by the extreme scarcity of fault samples, leading to suboptimal model performance and limited generalization. This work addresses this challenge by constructing a problem-agnostic hyperspherical synthetic dataset to systematically evaluate 14 anomaly detection algorithms—including kNN, LOF, XGBOD, SVM, and CatBoost—under rigorously controlled conditions across varying fault rates (0.05%–20%) and training set sizes. The study quantitatively demonstrates for the first time that unsupervised methods achieve optimal performance when fewer than 20 fault samples are available; semi-supervised and supervised approaches significantly outperform others with 30–50 fault samples; and further increasing normal samples yields diminishing returns. Additionally, feature dimensionality is found to critically influence the efficacy of semi-supervised methods, thereby clarifying the operational boundaries of each algorithmic category.

anomaly detectionclass imbalancefaulty data scarcity

This study addresses the challenge of estimating failure probabilities in complex systems under threshold conditions by proposing a novel approach that integrates Gabriel editing sets with a penalty-contour support vector machine. The method employs adaptive sampling to efficiently allocate points near the failure boundary, enabling the construction of geometrically consistent local linear surrogate models. This strategy significantly reduces the number of costly simulation calls while maintaining high accuracy, interpretability, and theoretical convergence guarantees. Empirical evaluations on four benchmark problems demonstrate superior performance compared to several state-of-the-art classifiers. Furthermore, the approach is successfully applied to estimate survival probabilities in the Lotka–Volterra competitive species model, confirming its effectiveness and practical utility in real-world scenarios.

computer modeldecision boundarymachine learning

Hot Scholars

SH

Steve Hanneke

Purdue University
Learning TheoryStatisticsArtificial Intelligence
SM

Shay Moran

(Math, CS, & DDS, Technion) & (Google Research)
Computer ScienceMathematics
MM

Mehryar Mohri

Head, ML Theory, Google Research; Professor, Courant Institute of Mathematical Sciences.
Machine Learning
TK

Tomer Koren

Associate Professor at Tel Aviv University
Machine LearningOptimizationReinforcement Learning
JA

Julian Asilis

Ph.D. Student, University of Southern California
Machine LearningLearning TheoryTheoretical CS