evaluation-to-data mapping

Designs and implements mapping schemas, taxonomies, and pipelines that convert evaluation outcomes (e.g., benchmark failures or capability slices) into concrete, targeted data interventions and test cases. Builds and analyzes closed-loop evaluation→data processes with rules and validation experiments to apply and verify data-level fixes while disambiguating data defects from model faults.

evaluation-to-datamapping

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.61
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.

capability slicedata-evaluation gapevaluation-to-data inference

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.

benchmarkingdata qualityevaluation framework

Automated Test Case Repair Using Language Models

Jan 12, 2024
AS
Ahmadreza Saboor Yaraghi
🏛️ University of Ottawa | Carleton University | University of Limerick

Software evolution frequently causes test failures, leading to high maintenance costs and low efficiency. This paper proposes TaRGet, the first framework to formalize test repair as a context-aware code translation task. Leveraging pre-trained models such as CodeT5 and CodeLlama, TaRGet automates repair via failure-context extraction, error-pattern-aware input construction, and two-stage fine-tuning. Its key contributions are: (1) a novel formalization of test repair as translation; (2) TaRBench—the first large-scale, empirically grounded benchmark comprising over 45K real-world test repairs; and (3) a reliability-prediction guidance mechanism, empirically validated for cross-project cold-start generalization. On TaRBench, TaRGet achieves a 66.1% exact-match repair rate—significantly outperforming state-of-the-art baselines—without requiring any project-specific training data.

Maintenance CostQuality AssuranceSoftware Testing

Latest Papers

What's happening recently
View more

Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.

auditable recordsevidence validationindustrial machine learning

Hot Scholars

TG

Tanya Goyal

Cornell University
Natural Language Processing
MM

Matteo Marsili

Senior reserch scientist, Abdus Salam ICTP, Trieste
Statistical mechanicsstochastic processescollective phenomena in socio-economic systemsnetworks
BC

Bhavya Chopra

University of California, Berkeley
Human-Computer InteractionData ScienceSoftware Engineering
EY

Emine Yilmaz

University College London
Information RetrievalNatural Language ProcessingMachine Learning