model quality

Designs and implements evaluation criteria, metrics, tests, and monitoring pipelines to measure model performance, calibration, robustness, fairness, and reliability. Uses error analysis, uncertainty quantification, validation procedures, and model selection/rollback processes to improve and maintain model quality across development and deployment.

modelquality

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.66
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$220K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.

Evaluation MetricsLimitationsMachine Learning Calibration

Technique to Baseline QE Artefact Generation Aligned to Quality Metrics

Nov 18, 2025
EF
Eitan Farchi
🏛️ IBM Research | IBM Consulting

This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.

Ensuring generated requirements and test cases meet quality metricsEstablishing baselines for automated QE artefact quality evaluationValidating LLM outputs through reverse generation and iterative refinement

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

Aug 25, 2025
JW
Jelke Wibbeke
🏛️ Carl von Ossietzky Universität Oldenburg | German Aerospace Center (DLR) | Jade University of Applied Science

In safety-critical applications, evaluating uncertainty calibration of regression models is hindered by inconsistent metric definitions, conflicting assumptions, and incomparable scales—impeding interpretability and reproducibility. This work systematically surveys and categorizes existing calibration metrics, then conducts a model-agnostic benchmark across real-world, synthetic, and manually miscalibrated datasets. We empirically demonstrate—for the first time—that most metrics yield contradictory or even opposing conclusions for identical calibration states, confirming that metric choice critically influences research outcomes. To address this, we propose ENCE (Expected Normalized Calibration Error) and CWC (Weighted Coverage Confidence) as more robust and stable primary metrics. Experiments across diverse scenarios show that ENCE and CWC exhibit superior consistency, strong resilience to noise and distribution shifts, and enhanced interpretability. Our findings establish a reproducible methodological foundation for uncertainty calibration evaluation in regression.

Evaluating conflicting calibration metrics for regression modelsIdentifying inconsistencies in recalibration metric performance comparisonsSystematically benchmarking reliability of uncertainty quantification methods

This work addresses a critical gap in the testing of machine learning (ML) components, where existing approaches predominantly focus on model performance while neglecting system-level quality attributes such as throughput, resource consumption, and robustness—often leading to integration failures. To bridge this gap, the paper proposes the first standalone quality model specifically tailored for ML components. Grounded in the ISO/IEC 25010 quality standard framework and informed by requirements engineering and software quality modeling techniques, the model systematically decouples and structures key quality attributes of ML components, thereby addressing the lack of component-level applicability in ISO/IEC 25059. It provides developers and stakeholders with a unified terminology to prioritize testing efforts. The model’s effectiveness has been validated through user studies and has been integrated into an open-source ML testing tool, enabling practical deployment.

ISO 25059machine learning componentsquality model

Quality in model-driven engineering: a tertiary study

Jun 23, 2016
MG
M. Goulão
🏛️ Universidade Nova de Lisboa | University of Maribor

Empirical evidence on the impact of Model-Driven Engineering (MDE) on software quality is fragmented and lacks systematic integration. Method: This paper conducts the first tertiary study dedicated to MDE quality research, systematically analyzing 22 published systematic literature reviews and mapping studies. It establishes a three-tier analytical framework to characterize research distribution, evidential strength, and methodological maturity in the MDE–quality domain. Results: Maintainability is the most studied quality attribute; however, among 83 identified research questions, 80 focus solely on conceptual or syntactic model-to-code mappings, with few conducting empirical comparisons. Crucially, MDE’s actual impact on quality in industrial development contexts remains markedly under-investigated. The study exposes a structural bias toward “re-modeling over validation” in current research and identifies critical gaps requiring urgent attention: rigorous experimental design, industry-based empirical validation, and multi-attribute quality assessment frameworks.

Aggregating consolidated findings on quality impact in model-driven engineeringAnalyzing software quality attributes most affected by MDE approachesIdentifying under-explored research areas needing further empirical validation

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study addresses the reliability of probabilistic uncertainty quantification (UQ) in software defect prediction, particularly its ability to reflect model performance and calibration—especially in cross-project settings, where systematic validation remains lacking. Through a large-scale empirical analysis of 16 classifiers across 36 within-project and 32 cross-project datasets, the work examines the relationships between five UQ metrics and six performance measures alongside three calibration metrics. It reveals, for the first time, a strong context dependency: within-project, UQ correlates strongly with false positive rate and AUC, but these correlations substantially weaken or even reverse in cross-project scenarios. Notably, high-performing models can still exhibit severe miscalibration. These findings indicate that UQ signals are not directly transferable and must be evaluated independently relative to specific objectives, using multidimensional calibration assessments.

CalibrationCross-Project PredictionPerformance Evaluation

This study addresses the widespread neglect in machine learning research of when validation occurs during data annotation—a critical factor influencing both label quality and cost—despite overreliance on post-hoc quality control. Drawing inspiration from the “shift-left” principle in software engineering, this work proposes a tripartite classification of quality checkpoints across early, intermediate, and late stages of the annotation pipeline and introduces a parameterized error propagation model that, for the first time, treats validation timing as a quantifiable design variable. Through error propagation modeling, process decomposition, and literature analysis, the authors find that only 4% of recent studies report validation timing. Their analysis demonstrates that early-stage quality checks can reduce error correction costs by up to two orders of magnitude. The paper calls for standardized reporting of timing configurations, platform support for tunable timing parameters, and empirical studies on stage-specific detection rates.

annotation pipelinesdata qualityerror propagation