supervised learning

Designs, implements, and evaluates models and training procedures that learn mappings from input features to labeled outputs using labeled datasets; this includes selecting model architectures, loss functions, optimization algorithms, feature representations, and regularization, running training and validation, and tuning hyperparameters. Analyzes model generalization, bias/variance tradeoffs, evaluation metrics and error sources, and prepares trained models for deployment or further experimentation.

supervisedlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Beyond algorithm hyperparameters: on preprocessing hyperparameters and associated pitfalls in machine learning applications

Dec 04, 2024
CS
Christina Sauer
🏛️ LMU Munich | Munich Center for Machine Learning | Medical University of Vienna

This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.

Addresses overlooked preprocessing hyperparameters in ML model tuningAims to improve predictive modeling quality and reportingHighlights pitfalls in informal preprocessing optimization practices

This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.

Addressing challenges in online, constrained, and multi-objective hyperparameter tuningAutomating hyperparameter search to improve machine learning efficiencyComparing state-of-the-art hyperparameter optimization techniques and methods

Three Mechanisms of Feature Learning in a Linear Network

Jan 13, 2024
YX
Yizhou Xu
🏛️ Abdus Salam International Center for Theoretical Physics | Massachusetts Institute of Technology | NTT Research

This work investigates how neural network width governs training dynamics. For single-hidden-layer linear networks, we derive the first exact analytical solution of learning dynamics at arbitrary finite width, unifying the characterization of the two-phase evolution—kernel learning and feature learning—and establishing a complete phase diagram parameterized by width, layer-wise learning rates, and initialization scale. Methodologically, we integrate analytical dynamical systems analysis, phase-diagram modeling, and empirical validation on nonlinear networks. Crucially, we identify three novel mechanisms operative during the feature-learning phase: alignment learning, de-alignment learning, and rescaling learning—each transcending the conventional kernel-method paradigm. These theoretical insights are empirically reproduced in realistic deep networks, offering a new conceptual framework for understanding training dynamics and designing adaptive optimization algorithms. (138 words)

Analyzes learning dynamics in neural networksExplores hyperparameter impact on training trajectoriesIdentifies feature learning mechanisms in networks

Learning Program Behavioral Models from Synthesized Input-Output Pairs

Jul 11, 2024
TM
Tural Mammadov
🏛️ CISPA Helmholtz Center for Information Security | Saarland University

This work addresses black-box program behavior modeling by proposing a reversible, differentiable, and constraint-aware, grammar-driven neural modeling framework. Methodologically, it generates input-output (I/O) pairs from formal grammars of input and output languages, and employs a lightweight (<6.3M-parameter) sequence-to-sequence model to cast program I/O mapping as a bidirectional neural machine translation task—enabling both forward prediction and backward inference—while supporting fine-grained behavioral constraints and fault- or coverage-guided input synthesis. Its key contribution is the first end-to-end joint modeling of reversibility, differentiability, and syntactic consistency in program behavior models. Evaluated on structured tasks such as Markdown and HTML generation, the framework achieves 95.4% accuracy and a BLEU score of 0.98±0.04, significantly outperforming existing irreversible or syntax-agnostic approaches.

Assists in program understanding and maintenance through synthesis.Learns program behavior models from input-output pairs.Predicts outputs and inputs using neural machine translation.

In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.

hyperparameter optimizationlarge-scale pre-traininglearning rate

Latest Papers

What's happening recently
View more

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Train on Validation (ToV): Fast data selection with applications to fine-tuning

Sep 30, 2025
AJ
Ayush Jain
🏛️ Granica Computing Inc. | Stanford University

Data selection for fine-tuning under scarce target-distribution samples remains challenging. Method: This paper proposes a “validation-set-driven data selection” paradigm: it swaps the conventional roles of validation set and training pool—performing lightweight fine-tuning on the validation set and selecting the most discriminative samples from the training pool based on the magnitude of prediction shifts induced by fine-tuning. The method requires no additional annotations or gradient computations, ensuring both efficiency and theoretical interpretability. Results: Evaluated on instruction tuning and named entity recognition, it significantly reduces test log-loss on the target distribution, consistently outperforming existing SOTA methods on average while improving data utilization efficiency and fine-tuning performance. Its core innovation lies in the first use of the validation set as a proxy for fine-tuning and leveraging prediction shift as the selection criterion—enabling precise identification of high-information samples under few-shot settings.

Improving data selection efficiency by reversing train-validation rolesReducing test loss through samples most affected by validation fine-tuningSelecting optimal training samples for fine-tuning with limited target data

The impact of training data quality on classifier performance is often overlooked. In the context of metagenomic DNA sequence assembly, this study systematically evaluates the behavior of Bayesian classifiers, neural networks, partition models, and random forests under various training data degradation scenarios. The findings reveal that as data quality deteriorates, all classifiers exhibit a “catastrophic” degradation pattern—shifting from substantially correct predictions to essentially random guesses. Concurrently, decision boundaries become sparser, and inter-classifier agreement paradoxically increases, indicating a convergence in error patterns under low-quality training conditions. This work provides the first quantitative characterization of the relationship between data quality and heterogeneity in classifier behavior, offering new insights for designing robust classification systems in data-scarce or noisy environments.

classifier congruenceclassifier performancedata degradation

Hot Scholars

GT

Giulio Turrisi

Researcher at the Dynamic Legged Systems Lab, Istituto Italiano di Tecnologia
roboticsmachine learningcontrolreinforcement learning
LL

Liang Lin

Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
CS

Claudio Semini

Head of the Dynamic Legged Systems Lab at Istituto Italiano di Tecnologia
roboticslocomotionquadrupedshydraulics
CZ

Chuxu Zhang

Associate Professor of CSE, University of Connecticut (UConn)
Machine LearningDeep LearningData Mining