Score
Designs and runs dataset-survey and benchmarking workflows that identify and quantify gaps between training and target data distributions—such as missing scenarios or labels, sensor or domain shifts, and synthetic-vs-real discrepancies—and produce numerical gap metrics and aligned dataset subsets. Builds analyses and visualizations to isolate dataset-driven versus model-/optimization-driven failure modes and to recommend concrete dataset collection, alignment, or benchmarking procedures.
This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
Large-scale, heterogeneous software defect datasets hinder efficient navigation and reuse by researchers. This paper systematically surveys 132 publicly available defect datasets and proposes a multidimensional evaluation framework—covering domain coverage, defect types, programming language distribution, construction methodologies, accessibility, and citation contexts—to achieve the first standardized metadata harmonization and empirical usability validation across a hundred-plus datasets. Through bibliometric analysis, citation network mapping, and cross-dimensional clustering, we identify test generation and automated program repair as the most widely supported application domains, while critical defect categories—including concurrency and security vulnerabilities—remain severely underrepresented. We further propose actionable dataset curation guidelines and a reusable assessment template, diagnosing pervasive issues such as incomplete coverage, disorganized structure, and outdated maintenance. Our work establishes a robust, empirically grounded data foundation for software defect detection, localization, repair, and AI-driven development.
Public datasets in the LLM4RE (Large Language Models for Requirements Engineering) domain are fragmented, poorly documented, and lack systematic description, hindering comparability and reuse. Method: We conduct the first systematic dataset mapping study in LLM4RE, analyzing 62 publicly available datasets drawn from 43 scholarly publications along dimensions including document type, granularity, RE task phase, domain, and language. We propose the first domain-specific dataset classification and characterization framework for LLM4RE. Contribution/Results: Our framework identifies critical research gaps—particularly in requirements elicitation, requirements management, and multilingual support—and we release an open dataset catalog alongside a standardized featureization schema. This work significantly enhances dataset visibility, structural consistency, and cross-study comparability, laying the foundation for a unified benchmarking repository in LLM4RE.
Dataset quality defects—such as missing documentation, incorrect labels, and ethical risks—are pervasive in open platforms yet resistant to detection by rule-based scripts, necessitating intelligent, automated identification methods. Method: We introduce the first LLM-agent benchmark for discovering real-world dataset quality issues, covering 221 empirically validated cases across eight platforms. It uniquely evaluates agents’ ability to autonomously detect latent defects without prior prompting. We propose an automated evaluation framework powered by GPT-4o, achieving high agreement with human experts (Cohen’s κ = 0.89), and ensure benchmark reliability via multi-source real-data sampling and expert annotation. Contribution/Results: Experiments reveal that even the state-of-the-art Curator agent detects only ~30% of defects, underscoring task difficulty. All benchmark data, code, and evaluation tools are publicly released to advance intelligent data governance.
Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.
This work addresses the limitation of existing AI benchmarks, which predominantly assess isolated data science capabilities while neglecting systematic evaluation of end-to-end project completion. The authors propose the first comprehensive evaluation framework tailored to full-cycle data science projects, introducing a benchmark comprising 40 real-world tasks that integrate multidimensional competencies—including technical implementation, analytical reasoning, communication, and ethical considerations. They further develop an assessment pipeline combining structured scoring rubrics with automated evaluation procedures. Experimental results demonstrate that state-of-the-art generative AI models perform comparably to junior data scientists on well-structured tasks, yet exhibit substantial performance gaps in tasks requiring subjective judgment, thereby underscoring the continued necessity of human validation in complex data science workflows.
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.
This study addresses the lack of systematic and impartial evaluation of existing methods for measuring distributional similarity in numerical data, which hinders informed selection in practice. The authors construct the first comprehensive benchmarking framework encompassing 36 similarity measures for continuous data—including statistical tests, distance-based metrics, and embedding approaches—and evaluate their discriminative power and computational efficiency through large-scale simulations across diverse distributional discrepancies (e.g., shifts in location, scale, and higher-order moments) and both two-sample and multi-sample settings. Based on empirical performance, the work proposes a data-characteristic-driven strategy for method selection, establishes a performance ranking, and demonstrates that combining only four to six methods suffices to achieve near-optimal performance in 90%–95% of scenarios.
This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.
This study addresses the limited reliability of existing training data contamination detection methods in real-world auditing scenarios, particularly when distribution shifts occur or when reference benchmarks are substantially smaller than the pretraining corpus. Through a systematic evaluation of three dominant paradigms—LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC—the authors conduct 335 experiments across 27 open-source and state-of-the-art closed-source language models (up to 27B parameters). They identify distribution shift and small-scale benchmarks as two critical failure modes, revealing that only 199 evaluations yield correct conclusions. Current approaches suffer from high false-positive rates, low statistical power, or coarse-grained provenance resolution, rendering them inadequate for reliably verifying individual benchmark subsets and underscoring the irreplaceable value of transparent data provenance.