Score
Designs and implements methods and pipelines that score, rank, filter, and sample data according to measured quality to produce higher‑quality training, validation, or evaluation subsets. Analyzes selection criteria and trade‑offs (e.g., quality thresholds, sampling strategies, and matching subset sizes) and evaluates how those choices affect downstream model performance and evaluation metrics.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This work addresses a critical gap in the testing of machine learning (ML) components, where existing approaches predominantly focus on model performance while neglecting system-level quality attributes such as throughput, resource consumption, and robustness—often leading to integration failures. To bridge this gap, the paper proposes the first standalone quality model specifically tailored for ML components. Grounded in the ISO/IEC 25010 quality standard framework and informed by requirements engineering and software quality modeling techniques, the model systematically decouples and structures key quality attributes of ML components, thereby addressing the lack of component-level applicability in ISO/IEC 25059. It provides developers and stakeholders with a unified terminology to prioritize testing efforts. The model’s effectiveness has been validated through user studies and has been integrated into an open-source ML testing tool, enabling practical deployment.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
NLP model evaluation faces critical challenges including benchmark saturation, data contamination, and imbalanced test instance quality. To address these, we propose SMART, a novel filtering framework that systematically integrates three criteria—triviality removal, contamination detection, and embedding-space diversity constraints—to automatically construct a compact, high-information, high-challenge, and low-redundancy subset from existing benchmarks. Specifically, SMART identifies trivial instances via prediction confidence thresholds, detects contamination through training-set overlap analysis, and enforces semantic diversity via clustering in pretrained embedding space. Evaluated on three multiple-choice QA benchmarks, SMART achieves an average 48% compression ratio. It attains significantly higher Pearson correlation with human evaluations (e.g., Chatbot Arena) than baseline benchmarks, while strictly preserving relative model rankings. This ensures enhanced evaluation reliability, computational efficiency, and robustness against dataset artifacts.
To address the inefficiency of foundation model training caused by high noise levels and label scarcity in internet-scale data, this paper proposes Mimic Score—a novel data quality metric. It leverages a pre-trained reference model to automatically assess the utility of individual samples for training new models by quantifying the alignment between their parameter-space gradients and those of the reference model. This approach pioneers gradient-direction alignment as the core criterion for data quality assessment, eliminating reliance on human annotations or downstream task validation. Based on Mimic Score, we introduce Grad-Mimic—an automated data filtering framework. Extensive experiments across six image datasets demonstrate that Grad-Mimic significantly improves model performance, particularly enhancing CLIP training efficacy, outperforming existing data curation methods, and enabling highly accurate estimation of dataset quality.
This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.
This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.
Existing approaches treat data quality assessment and machine learning systems as disjoint components, hindering dynamic, real-time coordination in production environments. This paper proposes the first end-to-end, quality-driven framework tailored for industrial MLOps, achieving the first closed-loop integration of data quality evaluation and model inference. The framework introduces theoretically grounded yet engineering-practical mechanisms: dynamic distribution drift detection, adaptive multi-dimensional quality metrics, a lightweight inference pipeline, and configurable quality thresholding. Evaluated on an industrial steelmaking ESR vacuum pump process, it achieves a model R² of 94%—a 12-percentage-point improvement—and reduces prediction latency by 75%, enabling millisecond-level quality-aware decision-making.
This study addresses the widespread neglect in machine learning research of when validation occurs during data annotation—a critical factor influencing both label quality and cost—despite overreliance on post-hoc quality control. Drawing inspiration from the “shift-left” principle in software engineering, this work proposes a tripartite classification of quality checkpoints across early, intermediate, and late stages of the annotation pipeline and introduces a parameterized error propagation model that, for the first time, treats validation timing as a quantifiable design variable. Through error propagation modeling, process decomposition, and literature analysis, the authors find that only 4% of recent studies report validation timing. Their analysis demonstrates that early-stage quality checks can reduce error correction costs by up to two orders of magnitude. The paper calls for standardized reporting of timing configurations, platform support for tunable timing parameters, and empirical studies on stage-specific detection rates.
This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.