benchmark classifiers

Designs and executes benchmarking experiments and evaluation pipelines to compare and rank classification models; this includes selecting datasets and splits, implementing training and evaluation code, computing performance metrics (e.g., F1, accuracy, ROC) and ensuring reproducible, fair comparisons. Analyzes results to identify best-performing classifiers, quantify statistical significance and variability, and diagnose model failure modes to inform model selection and improvement.

benchmarkclassifiers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.

benchmarkingdataset selectionmodel ranking

Machine Learning Pipeline for Software Engineering: A Systematic Literature Review

Jul 31, 2025
SK
Samah Kansab
🏛️ École de technologie supérieure (ÉTS)

Addressing the longstanding challenge in software engineering (SE) of balancing quality assurance with development efficiency, this study proposes a systematically derived and optimized machine learning (ML) pipeline framework tailored for SE tasks. Methodologically, the pipeline integrates automated data acquisition, SMOTE-based class balancing, SZZ-inspired feature selection, ensemble models (Random Forest and Gradient Boosting), and a novel evaluation metric—Balanced Accuracy Metric (BAM)—alongside standard metrics (AUC, F1-score, precision) and bootstrap resampling for robust validation. Key contributions include: (1) the first structured, end-to-end optimization framework for SE-specific ML pipelines; (2) empirical evidence demonstrating superior defect prediction performance of ensemble methods over individual classifiers; and (3) identification of two critical research gaps—insufficient data standardization and lack of model interpretability—thereby establishing foundational theoretical insights and actionable directions for intelligent SE research.

Addressing scalability issues in traditional SE approaches with MLEnsuring quality and efficiency in complex software engineering lifecycleOptimizing ML pipelines for robust defect prediction and automation

A Catalog of Fairness-Aware Practices in Machine Learning Engineering

Aug 29, 2024
GV
Gianmario Voria
🏛️ University of Salerno

The widespread deployment of machine learning (ML) in decision-making systems introduces significant fairness risks—particularly concerning the handling of sensitive attributes and the protection of minority groups—while software engineering lacks a systematic, lifecycle-oriented framework for fairness engineering practices. Method: We conduct a systematic mapping study (SMS) combined with a comprehensive literature review to analyze fairness-related practices across the ML development lifecycle. Contribution/Results: We propose the first software engineering–centric fairness practice taxonomy, comprising 28 structured, actionable practices explicitly mapped to data preprocessing, modeling, and deployment stages. Each practice is annotated with its corresponding ML lifecycle phase and contextual applicability, thereby bridging the gap between fairness research and industrial implementation. This taxonomy serves as an integrable, operational guide for researchers and practitioners, enhancing the reliability, accountability, and trustworthiness of ML systems.

Addressing fairness gaps in ML lifecycle practicesCataloging fairness-aware methods for sensitive feature treatmentProviding actionable fairness guidelines for ML engineering

Latest Papers

What's happening recently
View more

This study addresses the limitations of evaluating multiclass classifiers using single performance metrics, which often leads to misleading conclusions. To overcome this, the work proposes a multidimensional evaluation paradigm that leverages the PyCM library to construct a comprehensive analytical framework, enabling systematic comparison of classifier performance across a diverse set of evaluation metrics. Through two case studies, the research uncovers nuanced performance trade-offs that conventional metrics fail to capture, thereby demonstrating the necessity and effectiveness of multidimensional assessment in model selection and optimization. The findings further highlight the unique value of PyCM in facilitating thorough and precise evaluation of multiclass classification systems.

classifier comparisonevaluation frameworkmodel evaluation

This work addresses the limitation of existing AI benchmarks, which predominantly assess isolated data science capabilities while neglecting systematic evaluation of end-to-end project completion. The authors propose the first comprehensive evaluation framework tailored to full-cycle data science projects, introducing a benchmark comprising 40 real-world tasks that integrate multidimensional competencies—including technical implementation, analytical reasoning, communication, and ethical considerations. They further develop an assessment pipeline combining structured scoring rubrics with automated evaluation procedures. Experimental results demonstrate that state-of-the-art generative AI models perform comparably to junior data scientists on well-structured tasks, yet exhibit substantial performance gaps in tasks requiring subjective judgment, thereby underscoring the continued necessity of human validation in complex data science workflows.

AI benchmarkingautomated evaluationdata science workflow

This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.

empirical studyevaluation harnessesmachine learning

Existing evaluation benchmarks for large language models in software engineering often suffer from narrow task coverage, single-dimensional metrics, lack of realistic context, and data contamination, limiting their ability to comprehensively assess model robustness, fairness, and practical utility. To address these limitations, this work proposes BEHELM—the first full-stack benchmarking framework tailored for software engineering. BEHELM establishes a unified, standardized, and reproducible evaluation infrastructure through structured software scenario modeling, multi-granularity input-output specifications, and a multidimensional quality metric system encompassing robustness, explainability, fairness, and efficiency. By significantly lowering the barrier to constructing high-quality benchmarks, BEHELM enables systematic cross-task, cross-language, and cross-granularity evaluations, offering the community a more equitable, realistic, and future-oriented assessment paradigm.

benchmarkingdataset contaminationevaluation metrics

Hot Scholars

FB

Fadi Boutros

Research scientist, Fraunhofer Institute for Computer Graphics Research IGD
BiometricsFace recognitionGenerative AIComputer Vision
GL

Guang Li

Assistant Professor, Hokkaido University
Dataset DistillationSelf-Supervised LearningData-Centric AIMedical Image Analysis
KW

Kevin W. Bowyer

Schubmehl-Prein Family Professor of Computer Science and Engineering, University of Notre Dame
BiometricsPattern RecognitionComputer VisionData Mining
FY

François Yvon

ISIR / CNRS et Sorbonne Université
Natural Language ProcessingSpeech ProcessingComputational LinguisticsMachine Translation
ND

Naser Damer

Professor, TU Darmstadt and Fraunhofer Institute for Computer Graphics Research IGD
BiometricsFace recognitionComputer visionGenerative AI