dataset integration and alignment

Designs, builds, and evaluates reproducible data pipelines and artifacts that combine, clean, align, split, and augment multiple datasets to produce coherent integrated corpora and benchmarks (including multilingual, cross-domain, reasoning, and verification-focused collections). This work includes schema and label alignment, quality verification and verification-data distillation, expert-annotation integration, deterministic/steerable processing, faithful data selection, and creation of appropriate train/validation/test and evaluation splits.

datasetintegrationandalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing synthetic data generation tools often suffer from complex workflows, inconsistent standards, and limited cross-modal extensibility, hindering their ability to meet the data demands of large language models in specialized domains and low-resource languages. This work proposes a configuration-driven, end-to-end open-source framework that standardizes multi-source data synthesis through a unified and controllable paradigm. Featuring a highly modular architecture, the framework flexibly adapts to diverse tasks and supports high-quality data generation across multiple pathways, modalities, and languages. Integrated with both a graphical user interface and command-line utilities, it significantly lowers the barrier to entry for users. Empirical evaluations demonstrate that the framework effectively balances generation efficiency and data quality across various scenarios, thereby accelerating the practical deployment of synthetic data in model training pipelines.

data scarcitymultilingualmultimodal

An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science

Feb 23, 2025
QZ
Qiuhai Zeng
🏛️ Pennsylvania State University | Carnegie Mellon University | International Monetary Fund

Large language models (LLMs) exhibit poor code reproducibility and lack systematic evaluation in data science tasks. Method: We propose an Analyst-Reviewer dual-role framework—the first LLM evaluation paradigm explicitly designed for computational reproducibility. We formally define and quantify workflow sufficiency and completeness for reproducing functionally equivalent code; design two novel reproducibility-enhancing prompting strategies; and develop the first principle-driven, automated evaluation framework integrating rule-based validation and program analysis across three datasets and 1,032 tasks. Contribution/Results: Evaluating five state-of-the-art models, we find a strong correlation between reproducibility and accuracy; our prompting strategies significantly improve average reproducibility rates; and all code is publicly released.

Enhance transparency in LLM workflowsEvaluate reproducibility of LLMsIntroduce reproducibility-enhancing prompting strategies

Semantic drift in enterprise data pipelines—caused by multilingual transformations—decouples metadata from downstream data semantics, undermining reproducibility, governance, and performance of RAG and text-to-SQL applications. To address this, we propose a fine-grained schema lineage extraction method leveraging multilingual parsing, chain-of-thought prompting (optimized for 1.3B–32B models), and human-in-the-loop evaluation. We introduce SLiCE (Schema Lineage Composite Evaluation), the first benchmark framework tailored for multilingual script lineage, alongside a high-quality dataset of 1,700 real-world annotated samples. Experiments show that open-weight 32B models match GPT-4’s lineage accuracy under standard prompting, demonstrating cost-effective lineage extraction. Our core contributions are: (1) a systematic formalization of semantic-faithful lineage; (2) the first open, multilingual schema lineage benchmark with rigorous annotations; and (3) a lightweight, efficient extraction paradigm enabling scalable, accurate lineage inference.

Addressing semantic drift in data reproducibility and governanceEvaluating lineage quality with structural and semantic metricsExtracting fine-grained schema lineage from multilingual enterprise pipelines

UniGen: A Unified Framework for Textual Dataset Generation Using Large Language Models

Jun 27, 2024
SW
Siyuan Wu
🏛️ Huazhong University of Science and Technology | University of Notre Dame | University of Maryland | Microsoft Research | University of Wisconsin-Madison | Lehigh University

Existing LLM-based text data generation methods suffer from systematic limitations in generalizability, controllability, diversity, and factual fidelity. This paper introduces the first unified LLM framework for general-purpose text dataset generation. It innovatively integrates attribute-guided generation with group-level consistency verification to enhance diversity; combines code-executed mathematical evaluation and retrieval-augmented generation (RAG) to ensure label accuracy and factual consistency; and supports fine-grained, user-specified constraints. The framework unifies GPT-4 and Llama3 backbones with programmable label validation, attribute-conditioned control, and multi-stage collaborative verification. Experiments demonstrate substantial improvements in synthetic data quality—particularly in LLM evaluation benchmark construction and data augmentation tasks—yielding measurable gains in model reasoning performance and agent capabilities, and enabling dynamic, evolution-aware evaluation.

Addressing generalization and controllability in synthetic data generationEnhancing diversity and truthfulness in LLM-generated datasetsProviding customizable data generation for specific user requirements

CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation

Sep 03, 2024
IZ
Ingo Ziegler
🏛️ University of Copenhagen | Center for Information and Language Processing (CIS) | LMU Munich

Addressing the challenge of constructing high-quality, domain-specific annotated data—often costly and labor-intensive—this paper proposes a few-shot-driven synthetic data generation paradigm. Given only a small set of user-provided examples, the method retrieves semantically relevant real-world text from large-scale web corpora and leverages instruction-tuned large language models (LLMs) to automatically generate well-formatted, task-specific synthetic training data. It is the first approach to synergistically integrate corpus retrieval with LLM-based augmentation, enabling zero human annotation, domain adaptability, and efficient few-shot generalization. Empirical evaluation across biomedical, medical, and commonsense question answering (QA), as well as summarization tasks, demonstrates that models trained on the generated data achieve a 46-point preference score improvement over human-annotated baselines in summarization, while QA models match or surpass the performance of general-purpose foundation models.

Generates synthetic datasets for specialized tasks efficientlyOutperforms human-curated and other synthetic data methodsUses corpus retrieval and LLM augmentation for customization

Latest Papers

What's happening recently
View more

This work addresses a critical limitation in conventional large language model (LLM) evaluation, which treats benchmark datasets as homogeneous aggregates and overlooks the heterogeneity among samples in cognitive, linguistic, and task-related attributes. The authors propose a dataset-centric meta-evaluation framework that introduces fine-grained, sample-level annotations across five dimensions: cognitive demand, language quality, task characteristics, contextual dependency, and ethical safety. For the first time, this approach enables multidimensional auditing of widely used benchmarks such as MMLU and ARC. By allowing dynamic subset composition aligned with specific evaluation objectives, the framework uncovers the diversity obscured by aggregate accuracy metrics and establishes a composable evaluation paradigm tailored to targeted capabilities—such as reasoning depth or ethical sensitivity—thereby substantially enhancing the precision and interpretability of LLM assessments.

benchmark heterogeneitydataset introspectionevaluation bias

Existing benchmarks inadequately assess the quality of structured outputs from large language models in multi-source scenarios, as they focus either on structural compliance or value correctness within a single modality. This work proposes the first cross-modal, source-agnostic evaluation framework for structured generation, which uniformly converts inputs from text, images (OCR-processed PDFs), and audio (AMI meeting transcripts) into textual contexts, constrains model outputs via JSON Schema, and constructs a complex, realistic dataset through multi-hop question answering. Evaluation of 21 state-of-the-art models across seven metrics reveals near-perfect structural compliance but markedly lower value accuracy—83.0%, 67.2%, and 23.7% for text, image, and audio sources, respectively—highlighting that extracting structured information from long, multi-source contexts remains a significant challenge.

benchmarklarge language modelsmulti-source evaluation

This work addresses the lack of traceability in existing domain-specific fine-tuning approaches, which often leads to blind and inefficient data augmentation. The authors propose a “programming with data” paradigm that treats structured knowledge representations as a unified foundation for both training and evaluation, drawing an analogy to software development: training data serve as source code, model training as compilation, evaluation as unit testing, and data refinement as debugging. This framework enables precise, concept- and reasoning-chain–oriented model repair through structured knowledge extraction, test-driven data engineering, concept-level gap analysis, and diagnosis of broken reasoning chains. Validated across 16 disciplines, the approach significantly enhances model performance without compromising general capabilities, and the authors release an open-source knowledge base, evaluation suite, and training corpora to support reproducibility and further research.

data engineeringdomain specializationknowledge transfer

This study addresses the challenge of low-quality metadata that hinders dataset discoverability and reuse, particularly in the context of large language model (LLM)-generated descriptions lacking empirical guidance on context selection and its impact on quality. Building a literature-based framework for description quality assessment, the authors conduct systematic ablation experiments across 252 real-world CSV datasets. They uncover a previously unreported “table-structure penalty” phenomenon: relying solely on table structure significantly degrades narrative quality. While representative data samples aid semantic grounding, they do not improve overall human-rated quality. The work further reveals that different LLMs exhibit consistent descriptive styles. Through LLM-as-a-judge evaluations, semantic attribute analysis, and large-scale experimentation, the study offers key recommendations for LLM-assisted data publishing: concise, relevant context yields better results than redundant input, and table structure should be used cautiously as a basis for generation.

context ablationdata reusedataset description

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
KC

Kwanghee Choi

University of Texas at Austin
SpeechMachine LearningComputational Linguistics
KM

Kyoung Mu Lee

Professor, Department of Electrical and Computer Engineering, Seoul National University
Computer VisionMachine LearningArtificial Intelligence