evaluation methods

Designs and applies systematic procedures, metrics, and experimental protocols to measure, compare, and diagnose the performance, validity, reliability, and robustness of models, algorithms, systems, or interventions. Builds benchmarks, evaluation datasets, statistical analyses, and reporting practices to produce reproducible, interpretable results and to identify failure modes, confounders, and biases.

evaluationmethods

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$210K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Best Practices for Machine Learning Experimentation in Scientific Applications

Nov 26, 2025
UM
Umberto Michelucci
🏛️ Lucerne University of Applied Sciences and Arts | ZHAW - Zurich University of Applied Sciences

Scientific machine learning experiments often suffer from distorted performance evaluations due to poor experimental design and inconsistent documentation. To address this, we propose a principled framework for ML experimentation tailored to scientific research, encompassing data preprocessing, model selection, cross-validation, and reporting—emphasizing reproducibility, fair comparison, and transparency. Our key contributions include two novel quantitative metrics: the Logarithmic Overfitting Ratio (LOR) and Composite Overfitting Score (COS), which jointly characterize overfitting severity and instability across cross-validation folds. Complementing these, we introduce standardized preprocessing protocols, rigorously defined strong baselines, and modular visualization templates for diagnostic analysis. Empirical evaluation demonstrates that our framework substantially enhances experimental rigor, reproducibility, and result credibility in scientific ML. It further enables robust performance assessment and cross-study comparability, providing systematic support for establishing reliable benchmarks.

Addressing misleading conclusions from poor baselines and validation practicesEnsuring reproducibility and fair comparison in scientific ML experimentsProviding structured workflow for robust model evaluation in research

Current benchmarks for evaluating toxicity in large language models exhibit underappreciated systematic biases that may lead to the deployment of unsafe models. This work systematically investigates how variations in task formulation—such as text completion versus summarization—input data domains, and evaluated models interact with multiple toxicity metrics. It reveals, for the first time, that both task type and data domain significantly influence toxicity scores. Experiments demonstrate that existing benchmarks are prone to misclassifying content as harmful when tasks are altered and show inconsistent performance across domains, highlighting their fragility and dependence on specific model-task configurations. These findings underscore the urgent need for more robust and reliable toxicity evaluation frameworks.

benchmark robustnessevaluation biasLLM evaluation

Experimental reproducibility in Empirical Software Engineering (ESE) is hindered by a fundamental disconnect between idealized methodological assumptions—e.g., standardized protocols and controlled conditions—and researchers’ actual experimental practices. Method: We conducted a two-year ethnographic study involving participant observation, in-depth interviews, and content analysis of experimental artifacts across diverse ESE research teams. Contribution/Results: We identify four critical dimensions—activity diversity, role distribution, conceptual granularity, and domain perspective—in which real-world experimentation systematically deviates from textbook models. Based on these findings, we propose the first high-fidelity conceptual and process model grounded in empirical research practice, explicitly capturing the “practice gap” underlying irreproducibility. This model provides foundational evidence and design principles for developing next-generation reproducibility-support tools, methodological guidelines, and evaluation frameworks in ESE.

Compares actual experimental processes with textbook methodologies in detailExplores mismatches between proposed replication procedures and researchers' needsInvestigates how experimental researchers conduct experiments in practice

This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.

analysis reasoningassumptionsdata analysis

Latest Papers

What's happening recently
View more

This study addresses a critical gap in existing scientific data analysis benchmarks, which fail to differentiate models’ capabilities across distinct scientific reasoning tasks—such as hypothesis exploration, causal inference, and mechanistic explanation. To this end, the authors introduce SDABench, the first multidimensional evaluation benchmark specifically designed to assess scientific analytical competence. It encompasses six dimensions: descriptive, exploratory, inferential, predictive, causal, and mechanistic reasoning, comprising 527 real-world and 6,000 synthetically generated data instances across five scientific domains. Using a five-stage error analysis framework, the benchmark systematically evaluates 15 prominent large language models. Results reveal strong performance on descriptive tasks but substantial deficiencies in complex reasoning involving hypothesis selection, latent variable modeling, and mechanistic inference, indicating that current models remain ill-equipped to support high-level scientific discovery.

capability-oriented benchmarklarge language modelsmechanistic reasoning

This study addresses the lack of executable and verifiable knowledge representations in existing meta-analyses, which hinders the traceability and reproducibility of critical analytical decisions. To overcome this limitation, the authors propose Executable Analytical Knowledge Representation (EAKR) and introduce MetaSynDec, an agent-based framework that, for the first time, enables explicit modeling, machine-actionable execution, and closed-loop validation of meta-analytic decisions. The system leverages large language models to generate structured knowledge and validates and executes it through deterministic, schema- and contract-based services. Evaluated across 58 synthesis units, EAKR successfully constructed all units, achieved exact evidence-set consistency in 75% of cases, and produced confidence intervals overlapping with published results in 98.2% of cases—substantially outperforming direct LLM-generated approaches.

analytical knowledge representationevidence synthesisexecutable knowledge

Hot Scholars

DM

Dinesh Manocha

Distinguished University Professor, University of Maryland at College Park
computer graphicsgeometric modelingmotion planningvirtual reality
KK

Kevin Klyman

Stanford, Harvard
Foundation ModelsAI RegulationGeopolitics
AA

Aishwarya Agrawal

University of Montreal, Mila, Google DeepMind
Artificial IntelligenceMultimodal Vision-LanguageComputer VisionNLP
YL

Yong Luo

Wuhan University
Artifical IntelligenceMachine LearningData MiningPattern Classification and Search
AR

Anka Reuel

CS Ph.D. Candidate, Stanford University
AI GovernanceResponsible AIAI EthicsAI Safety