ml research

Designs, implements, and evaluates new machine learning algorithms, model architectures, training objectives, and optimization procedures; and builds experiments, datasets, benchmarks, and reproducible code to test, analyze, and validate hypotheses about model performance, generalization, robustness, and theoretical properties.

mlresearch

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.82
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Best Practices for Machine Learning Experimentation in Scientific Applications

Nov 26, 2025
UM
Umberto Michelucci
🏛️ Lucerne University of Applied Sciences and Arts | ZHAW - Zurich University of Applied Sciences

Scientific machine learning experiments often suffer from distorted performance evaluations due to poor experimental design and inconsistent documentation. To address this, we propose a principled framework for ML experimentation tailored to scientific research, encompassing data preprocessing, model selection, cross-validation, and reporting—emphasizing reproducibility, fair comparison, and transparency. Our key contributions include two novel quantitative metrics: the Logarithmic Overfitting Ratio (LOR) and Composite Overfitting Score (COS), which jointly characterize overfitting severity and instability across cross-validation folds. Complementing these, we introduce standardized preprocessing protocols, rigorously defined strong baselines, and modular visualization templates for diagnostic analysis. Empirical evaluation demonstrates that our framework substantially enhances experimental rigor, reproducibility, and result credibility in scientific ML. It further enables robust performance assessment and cross-study comparability, providing systematic support for establishing reliable benchmarks.

Addressing misleading conclusions from poor baselines and validation practicesEnsuring reproducibility and fair comparison in scientific ML experimentsProviding structured workflow for robust model evaluation in research

This work proposes the first end-to-end automated artificial intelligence research framework capable of fully automating the development pipeline from algorithmic idea generation to executable machine learning classifiers. The approach integrates structured meta-prompt engineering with large language model–based code generation, augmented by an automated evaluation and iterative optimization mechanism. Experimental results on twenty standard datasets from the Infinity-Bench benchmark demonstrate that multiple novel classifiers autonomously generated by the framework significantly outperform baseline methods implemented in scikit-learn. This study thus achieves, for the first time, complete automation of the entire workflow—from initial algorithmic conception to deployable, runnable code—marking a significant step toward self-driving AI research systems.

AI automationautomate AI researchend-to-end framework

ExeKGLib: A Platform for Machine Learning Analytics based on Knowledge Graphs

Aug 01, 2025
AK
Antonis Klironomos
🏛️ Bosch Center for AI | Oslo Metropolitan University | RWTH Aachen | University of Mannheim | University of Oslo

To address the challenge that domain experts—lacking machine learning (ML) expertise—struggle to construct high-quality analytical pipelines, this paper proposes a knowledge graph–based low-code ML platform. Methodologically, it encodes ML best practices, algorithmic constraints, and domain semantics into a structured, inferable knowledge graph, enabling visual pipeline orchestration and automated execution via a graphical user interface. Technically, the platform integrates the Python ecosystem, modern GUI frameworks, and a robust pipeline automation engine. Its key contribution lies in being the first to deeply embed a reasoning-capable knowledge graph into ML workflow design, thereby significantly enhancing pipeline executability, transparency, and cross-domain reusability. Empirical validation across multiple real-world scientific and engineering case studies demonstrates the platform’s effectiveness: non-ML practitioners can independently build, interpret, and reuse high-quality analytical workflows.

Enables non-ML experts to build ML pipelines easilyImproves transparency and reusability of ML workflowsUses knowledge graphs to simplify ML pipeline creation

Existing AI agents lack systematic evaluation for Machine Learning Engineering (MLE) capabilities. Method: We introduce MLE-bench, the first benchmark dedicated to MLE competence—comprising 75 real-world Kaggle competition tasks spanning core engineering stages including data preprocessing, model training, and experiment management. We formally define and quantify MLE capability dimensions, establish a human baseline from Kaggle participants, and release an open-source, reproducible automated evaluation framework. To ensure assessment integrity, we integrate state-of-the-art open-agent frameworks (e.g., AIDE) with large language models (e.g., o1-preview) and conduct rigorous data contamination analysis. Results: The optimal configuration (o1-preview + AIDE) achieves Kaggle Bronze-level performance on 16.9% of tasks. Our analysis reveals critical dependencies of MLE capability on computational scaling and pretraining data contamination, providing foundational insights for agent development in ML engineering.

Assess performance on Kaggle competitionsEvaluate AI agents' ML engineering skillsInvestigate AI resource scaling impact

Conformal prediction under feedback covariate shift for biomolecular design

Feb 08, 2022
CF
Clara Fannjiang
🏛️ University of California, Berkeley

This work addresses the challenge of quantifying prediction uncertainty in generative biomolecular design, where feedback covariate shift undermines conventional uncertainty estimation. We propose the first conformal prediction framework tailored to closed-loop design paradigms. Departing from standard i.i.d. assumptions, our method imposes no structural constraints on either the design algorithm or the regression model, delivering finite-sample statistically valid confidence sets for arbitrary black-box design pipelines. Key innovations include quantile-regression-driven adaptive conformal prediction, explicit modeling of feedback-induced distributional shift, and robust error calibration. Evaluated on protein and small-molecule design tasks, our approach achieves ≥94.8% empirical coverage at the 95% nominal confidence level—substantially outperforming standard conformal methods (which drop to as low as 72%)—while maintaining high predictive accuracy.

Address distribution shift in training-test data dependenceConstruct confidence sets for model predictionsQuantify uncertainty in protein fitness predictions

Latest Papers

What's happening recently
View more

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.

AI-enabled systemsarchitectural designmachine learning integration

This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.

computational advertisinglarge-scale ML ecosystemsML technique propagation

This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.

empirical studyevaluation harnessesmachine learning

Existing AI research agents often produce seemingly plausible but ineffective machine learning solutions due to a lack of systematic training. To address this, this work proposes the first scalable synthetic task generation framework that automatically constructs high-quality, executable research tasks through topic sampling, proposal generation grounded in real-world Hugging Face datasets, and self-debugging validation. The framework further leverages trajectory distillation—transferring effective research behaviors from GPT-5 to Qwen3—to guide student models in learning valid scientific reasoning paths. Evaluated on the MLGym benchmark, Qwen3-4B and Qwen3-8B models trained with this approach achieve 9% and 12% relative improvements in Area Under the Performance curve (AUP), respectively, substantially outperforming baseline methods.

Agent TrainingAI ScientistAutomatic Scientific Discovery

Hot Scholars

KK

Kaan Kale

Undergraduate Student, Bogazici University
AT

Alexander Tessier

Autodesk Research, University of Toronto
Computer GraphicsSimulationVisualizationEngineering
FW

FuTe Wong

University of Toronto
Deep Learning/Quantum Machine Learning/Brain Stimulation
KM

Kyle Mylonakis

Protopia AI
Deep LearningPrivacyApplied Mathematics