applied machine learning

Designs, builds, and evaluates practical machine learning solutions and production-ready pipelines, covering data ingestion, preprocessing, feature engineering, model selection, training, tuning, validation, and deployment. Optimizes and instruments models for real-world constraints (scalability, latency, robustness, reproducibility and monitoring), and analyzes model performance and behavior against operational metrics and data-shift scenarios.

appliedmachinelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Is Your Training Pipeline Production-Ready? A Case Study in the Healthcare Domain

Jun 07, 2025
DL
Daniel Lawand
🏛️ University of São Paulo | Tilburg University | Technical University of Eindhoven

Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.

Ensuring ML training pipelines are production-ready in healthcareEvolving architecture for better maintainability and robustnessImproving software quality in MLES for respiratory pre-diagnosis

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Beyond algorithm hyperparameters: on preprocessing hyperparameters and associated pitfalls in machine learning applications

Dec 04, 2024
CS
Christina Sauer
🏛️ LMU Munich | Munich Center for Machine Learning | Medical University of Vienna

This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.

Addresses overlooked preprocessing hyperparameters in ML model tuningAims to improve predictive modeling quality and reportingHighlights pitfalls in informal preprocessing optimization practices

Machine learning model selection lacks formalized methodologies, making it difficult to systematically characterize contextual factors—such as data characteristics and prediction tasks—and their interactions, resulting in opaque, non-adaptive decisions. This paper introduces, for the first time, software product line (SPL) principles into ML model selection, proposing a variability-aware algorithm selection framework. It constructs a configurable feature model that explicitly captures commonalities and variabilities among contextual factors—including dataset size, feature dimensionality, and task type—as well as their logical dependencies. By integrating scikit-learn’s heuristic rules with an instantiation framework, the approach enables interpretable, adaptive, and transparent model recommendations. An empirical case study demonstrates that the method significantly outperforms existing strategies in accuracy, interpretability, and contextual adaptability.

Machine LearningModel SelectionRule Formalization

MLOps Monitoring at Scale for Digital Platforms

Apr 23, 2025
YJ
Yu Jeffrey Hu
🏛️ Purdue University | Essec Business School | Maastricht University

Massive, dynamic data streams in digital platforms render conventional ML monitoring methods ineffective or prohibitively costly in manual effort, forcing enterprises to downgrade to simpler models. Method: This paper proposes the Machine Learning Monitoring Agent (MLMA) framework, introducing a test-driven, automated retraining mechanism based on data-adaptive reference loss batches—designed to enable efficient closed-loop operations while preserving human-in-the-loop collaborative governance. The approach integrates design science principles, dynamic reference loss computation, key metric visualization, and human–AI collaborative workflows. Contribution/Results: Evaluated on a large-scale instant-delivery platform, MLMA supports concurrent monitoring of hundreds of models, significantly reduces manual intervention frequency, and sustains long-term online model performance stability. Its core contribution lies in unifying dynamic data adaptation, automated trigger logic, and human–AI collaboration—thereby overcoming critical technical bottlenecks in real-time monitoring and adaptive maintenance of large-scale ML systems.

Automating re-training to maintain model performance at scaleMonitoring ML models in large unstable data streamsReducing labor-intensive MLOps supervision in digital platforms

Latest Papers

What's happening recently
View more

This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.

capacity planningload testingML model serving

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This work proposes the first end-to-end automated artificial intelligence research framework capable of fully automating the development pipeline from algorithmic idea generation to executable machine learning classifiers. The approach integrates structured meta-prompt engineering with large language model–based code generation, augmented by an automated evaluation and iterative optimization mechanism. Experimental results on twenty standard datasets from the Infinity-Bench benchmark demonstrate that multiple novel classifiers autonomously generated by the framework significantly outperform baseline methods implemented in scikit-learn. This study thus achieves, for the first time, complete automation of the entire workflow—from initial algorithmic conception to deployable, runnable code—marking a significant step toward self-driving AI research systems.

AI automationautomate AI researchend-to-end framework

This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.

empirical studyevaluation harnessesmachine learning

This study addresses the critical yet underexamined role of data filtering in clinical machine learning, which alters statistical structures and directly impacts task complexity and model performance. Despite these effects, existing research frequently treats filtering as routine preprocessing with insufficient transparency. This work reconceptualizes data filtering as a core component of the scientific method, advocating its integration into the broader research paradigm rather than its treatment as a mere technical step. To this end, we develop a transparent and interpretable clinical data preprocessing pipeline and release the corresponding code as open source. Our analysis elucidates the mechanisms through which filtering decisions critically influence data distributions and downstream model efficacy. Ultimately, this research provides a novel framework for enhancing methodological rigor and reproducibility in clinical artificial intelligence studies.

Clinical Machine LearningData FilteringExplainability

Hot Scholars

SS

Sho Sonoda

RIKEN Center for Advanced Intelligence Project (AIP)
machine learningharmonic analysis
ZY

Zihan Yu

MSc Student at Imperial College London
Computer VisionMedical AI
JL

Jiazhen Liu

Georgia Insitute of Technology
RoboticsRobot LearningTask and Motion Planning
LE

Lisa Ehrlinger

Hasso Plattner Institute, University of Potsdam
Data QualityKnowledge GraphsSemantic TechnologyMetadata Management
JS

John Stufken

George Mason University
Design and Analysis of ExperimentsBig Data