Score
Designs, builds, and evaluates reproducible data pipelines and artifacts that combine, clean, align, split, and augment multiple datasets to produce coherent integrated corpora and benchmarks (including multilingual, cross-domain, reasoning, and verification-focused collections). This work includes schema and label alignment, quality verification and verification-data distillation, expert-annotation integration, deterministic/steerable processing, faithful data selection, and creation of appropriate train/validation/test and evaluation splits.
This paper addresses the insufficient bidirectional synergy between large language models (LLMs) and data management (DATA). To bridge this gap, it proposes a dual-path integration framework: DATA4LLM—establishing a data supply infrastructure for the full LLM lifecycle, incorporating techniques such as deduplication, synthetic data augmentation, KV cache optimization, RAG post-processing, and retrieval-augmented prompting; and LLM4DATA—leveraging LLMs as general-purpose data engines for novel paradigms in data manipulation, analysis, and system optimization. The work introduces the first systematic taxonomy spanning both LLM and database research domains, identifying 12 categories of technical challenges and surveying 87 state-of-the-art studies. This taxonomy provides a foundational theoretical framework and practical guidelines for designing AI-native data systems.
Existing synthetic data generation tools often suffer from complex workflows, inconsistent standards, and limited cross-modal extensibility, hindering their ability to meet the data demands of large language models in specialized domains and low-resource languages. This work proposes a configuration-driven, end-to-end open-source framework that standardizes multi-source data synthesis through a unified and controllable paradigm. Featuring a highly modular architecture, the framework flexibly adapts to diverse tasks and supports high-quality data generation across multiple pathways, modalities, and languages. Integrated with both a graphical user interface and command-line utilities, it significantly lowers the barrier to entry for users. Empirical evaluations demonstrate that the framework effectively balances generation efficiency and data quality across various scenarios, thereby accelerating the practical deployment of synthetic data in model training pipelines.
Large language models (LLMs) exhibit poor code reproducibility and lack systematic evaluation in data science tasks. Method: We propose an Analyst-Reviewer dual-role framework—the first LLM evaluation paradigm explicitly designed for computational reproducibility. We formally define and quantify workflow sufficiency and completeness for reproducing functionally equivalent code; design two novel reproducibility-enhancing prompting strategies; and develop the first principle-driven, automated evaluation framework integrating rule-based validation and program analysis across three datasets and 1,032 tasks. Contribution/Results: Evaluating five state-of-the-art models, we find a strong correlation between reproducibility and accuracy; our prompting strategies significantly improve average reproducibility rates; and all code is publicly released.
Semantic drift in enterprise data pipelines—caused by multilingual transformations—decouples metadata from downstream data semantics, undermining reproducibility, governance, and performance of RAG and text-to-SQL applications. To address this, we propose a fine-grained schema lineage extraction method leveraging multilingual parsing, chain-of-thought prompting (optimized for 1.3B–32B models), and human-in-the-loop evaluation. We introduce SLiCE (Schema Lineage Composite Evaluation), the first benchmark framework tailored for multilingual script lineage, alongside a high-quality dataset of 1,700 real-world annotated samples. Experiments show that open-weight 32B models match GPT-4’s lineage accuracy under standard prompting, demonstrating cost-effective lineage extraction. Our core contributions are: (1) a systematic formalization of semantic-faithful lineage; (2) the first open, multilingual schema lineage benchmark with rigorous annotations; and (3) a lightweight, efficient extraction paradigm enabling scalable, accurate lineage inference.
Existing LLM-based text data generation methods suffer from systematic limitations in generalizability, controllability, diversity, and factual fidelity. This paper introduces the first unified LLM framework for general-purpose text dataset generation. It innovatively integrates attribute-guided generation with group-level consistency verification to enhance diversity; combines code-executed mathematical evaluation and retrieval-augmented generation (RAG) to ensure label accuracy and factual consistency; and supports fine-grained, user-specified constraints. The framework unifies GPT-4 and Llama3 backbones with programmable label validation, attribute-conditioned control, and multi-stage collaborative verification. Experiments demonstrate substantial improvements in synthetic data quality—particularly in LLM evaluation benchmark construction and data augmentation tasks—yielding measurable gains in model reasoning performance and agent capabilities, and enabling dynamic, evolution-aware evaluation.
Addressing the challenge of constructing high-quality, domain-specific annotated data—often costly and labor-intensive—this paper proposes a few-shot-driven synthetic data generation paradigm. Given only a small set of user-provided examples, the method retrieves semantically relevant real-world text from large-scale web corpora and leverages instruction-tuned large language models (LLMs) to automatically generate well-formatted, task-specific synthetic training data. It is the first approach to synergistically integrate corpus retrieval with LLM-based augmentation, enabling zero human annotation, domain adaptability, and efficient few-shot generalization. Empirical evaluation across biomedical, medical, and commonsense question answering (QA), as well as summarization tasks, demonstrates that models trained on the generated data achieve a 46-point preference score improvement over human-annotated baselines in summarization, while QA models match or surpass the performance of general-purpose foundation models.
This work addresses a critical limitation in conventional large language model (LLM) evaluation, which treats benchmark datasets as homogeneous aggregates and overlooks the heterogeneity among samples in cognitive, linguistic, and task-related attributes. The authors propose a dataset-centric meta-evaluation framework that introduces fine-grained, sample-level annotations across five dimensions: cognitive demand, language quality, task characteristics, contextual dependency, and ethical safety. For the first time, this approach enables multidimensional auditing of widely used benchmarks such as MMLU and ARC. By allowing dynamic subset composition aligned with specific evaluation objectives, the framework uncovers the diversity obscured by aggregate accuracy metrics and establishes a composable evaluation paradigm tailored to targeted capabilities—such as reasoning depth or ethical sensitivity—thereby substantially enhancing the precision and interpretability of LLM assessments.
Existing benchmarks inadequately assess the quality of structured outputs from large language models in multi-source scenarios, as they focus either on structural compliance or value correctness within a single modality. This work proposes the first cross-modal, source-agnostic evaluation framework for structured generation, which uniformly converts inputs from text, images (OCR-processed PDFs), and audio (AMI meeting transcripts) into textual contexts, constrains model outputs via JSON Schema, and constructs a complex, realistic dataset through multi-hop question answering. Evaluation of 21 state-of-the-art models across seven metrics reveals near-perfect structural compliance but markedly lower value accuracy—83.0%, 67.2%, and 23.7% for text, image, and audio sources, respectively—highlighting that extracting structured information from long, multi-source contexts remains a significant challenge.
This work addresses the lack of traceability in existing domain-specific fine-tuning approaches, which often leads to blind and inefficient data augmentation. The authors propose a “programming with data” paradigm that treats structured knowledge representations as a unified foundation for both training and evaluation, drawing an analogy to software development: training data serve as source code, model training as compilation, evaluation as unit testing, and data refinement as debugging. This framework enables precise, concept- and reasoning-chain–oriented model repair through structured knowledge extraction, test-driven data engineering, concept-level gap analysis, and diagnosis of broken reasoning chains. Validated across 16 disciplines, the approach significantly enhances model performance without compromising general capabilities, and the authors release an open-source knowledge base, evaluation suite, and training corpora to support reproducibility and further research.
This study addresses the challenge of low-quality metadata that hinders dataset discoverability and reuse, particularly in the context of large language model (LLM)-generated descriptions lacking empirical guidance on context selection and its impact on quality. Building a literature-based framework for description quality assessment, the authors conduct systematic ablation experiments across 252 real-world CSV datasets. They uncover a previously unreported “table-structure penalty” phenomenon: relying solely on table structure significantly degrades narrative quality. While representative data samples aid semantic grounding, they do not improve overall human-rated quality. The work further reveals that different LLMs exhibit consistent descriptive styles. Through LLM-as-a-judge evaluations, semantic attribute analysis, and large-scale experimentation, the study offers key recommendations for LLM-assisted data publishing: concise, relevant context yields better results than redundant input, and table structure should be used cautiously as a basis for generation.