Score
Designs and implements pipelines, prompt templates, and validation tooling that use large language models to extract and structure factual knowledge or priors from unstructured inputs, producing typed and labeled facts. Work covers per-unit processing (for example, per-function), deriving dataflow facts, variable types and invariants, and serializing or formatting those facts for logical ingestion and downstream analysis (e.g., datalog), plus methods to curate and validate the extracted knowledge.
Existing automated data preparation tools lack robust semantic understanding and struggle with complex, context-dependent data quality issues. Method: This study investigates the efficacy of large language models (LLMs) in data profiling and cleaning on low-quality datasets. We propose a customized data quality assessment framework informed by a practitioner-focused user study, and systematically evaluate both general-purpose and fine-tuned table-centric LLMs—via prompt engineering—on tasks including anomaly detection, cleaning logic generation, and error repair, benchmarking against traditional tools (e.g., Trifacta, OpenRefine). Contribution/Results: LLMs significantly outperform conventional tools in contextual reasoning and generating interpretable, human-verifiable cleaning rules; however, their output precision and deterministic verifiability remain limited. This work establishes the first evaluation paradigm specifically designed for LLMs in data preparation and empirically validates their viability—and practical boundaries—as collaborative “data engineering partners.”
This work addresses the challenges of accuracy and scalability in knowledge graph fact verification at scale, where existing automated methods remain immature. We propose the first multidimensional evaluation framework for assessing large language models (LLMs) in this context, systematically examining their capabilities along three dimensions: internal knowledge, retrieval-augmented generation (RAG), and multi-model consensus. To support this evaluation, we construct a RAG dataset comprising two million documents, a FactCheck benchmark, and an interactive analysis platform, conducting experiments across three real-world knowledge graphs. Our results reveal that while LLMs show promise, their performance lacks robustness; furthermore, the effectiveness of RAG and multi-model strategies varies significantly, underscoring both the necessity of systematic evaluation and the practical utility of our proposed framework.
This paper addresses the factual inconsistency gap in fine-tuned large language models (LLMs) between known (in-distribution) and unknown (out-of-distribution) knowledge. Methodologically, it establishes— for the first time at the theoretical level—that test-time prompting techniques (e.g., in-context learning [ICL] and chain-of-thought [CoT]) can attenuate or even override the influence of fine-tuning data; it further introduces the “prompt dominance” theory, redefining evaluation criteria for fine-tuning data. Empirical results demonstrate that ICL and CoT improve factual accuracy by over 35% on out-of-distribution knowledge tasks. The core contribution lies in uncovering the compensatory mechanism whereby prompt engineering mitigates fine-tuning biases, thereby establishing a novel paradigm—“prompting to compensate for fine-tuning deficiencies.” This work provides interpretable, quantifiable theoretical principles and practical guidelines for synergistically optimizing fine-tuning and reasoning.
Current LLM-native software engineering lacks a systematic practical framework—particularly in verification and falsification—necessitating unified task taxonomies and prompt-pattern conceptualizations. Method: We conduct a systematic literature review of over 100 papers, employing bibliometric analysis and conceptual clustering to map, classify, and abstract LLM-based downstream tasks in software engineering (SE). Contribution/Results: We propose the first fine-grained SE-specific taxonomy for LLM downstream tasks, encompassing six core clusters: testing, fuzzing, bug localization, vulnerability detection, static analysis, and program verification. Our taxonomy uniquely balances cross-task abstraction with task-specific variation modeling, uncovering generalizable prompt-engineering principles. It provides a foundational framework for targeted LLM adaptation, benchmark construction, and empirically grounded engineering practice in SE.
Large language models (LLMs) frequently fail in real-world tool invocation due to intent misinterpretation, incorrect parsing of tool documentation, and parameterization errors. To address this, we propose a curriculum-inspired structured reasoning framework that replaces free-form chain-of-thought prompting with guided, template-based reasoning—explicitly decoupling the process into three sequential stages: *intent parsing*, *tool matching*, and *parameter generation*. Our framework employs stepwise structured prompts to jointly model user goals and tool functionalities, thereby enhancing invocation robustness and decision interpretability. Evaluated across multiple state-of-the-art models (e.g., LLaMA-3, Qwen2) and benchmarks (ToolBench, API-Bank), it reduces relative error rates by 3–12% over strong baselines. The core contribution lies in transforming implicit, unstructured reasoning into an explicit, traceable, and modular pipeline—balancing accuracy with transparency and auditability.
This work addresses the challenges of scarce labeled data and weakly expressed, low-salience product attributes in applications such as digital product passports by proposing a two-step verification generative information extraction framework that integrates pretrained language models (PLMs) with large language models (LLMs). The approach first employs a PLM for initial candidate extraction and then leverages a locally deployable open-source LLM—such as those in the Llama family—for secondary verification and error correction, substantially improving extraction accuracy for sparse and weakly expressed entities. Experimental results demonstrate that the proposed framework enhances generalization capability and enables medium-scale models to approach the performance of much larger models, all while preserving data privacy and maintaining computational efficiency. The method has been successfully integrated into a demonstration system tailored for digital product passports.
This study addresses the lack of systematic evaluation for Datalog programs generated by large language models (LLMs) by constructing a benchmark comprising 136 tasks and proposing reliable metrics based on execution verification and mutation analysis. Through experiments involving six prompting strategies applied to six LLMs and coding agents, results indicate that direct prompting achieves a maximum match rate of 68.4%, whereas coding agents attain 83.8% while effectively eliminating most compilation errors. The research precisely identifies semantic and compilation errors, revealing that recursive reasoning remains a core challenge for current models. Overall, this work establishes a new paradigm for evaluating the logical programming capabilities of LLMs.
The era of large language models (LLMs) faces critical challenges including insufficient high-quality data supply, fragmented data preparation pipelines, poor reproducibility, and lack of model-in-the-loop support. Method: We propose the first LLM-driven, unified data preparation framework for data-centric AI, featuring system-level abstractions and PyTorch-style APIs for modular design. We introduce DataFlow-Agent—the first agent that synthesizes executable data pipelines end-to-end from natural language specifications—and integrate LLM-powered operator synthesis, iterative validation, 200+ reusable operators, and six domain-agnostic pipeline templates. Results: Experiments on Text-to-SQL, code generation, and mathematical reasoning show our synthesized data significantly outperforms human-annotated and domain-specific synthetic data. Remarkably, just 10K samples surpass the performance of models trained on the million-scale Infinity-Instruct dataset, empirically validating the decisive impact of data quality on model performance.
This work addresses the lack of traceability in existing domain-specific fine-tuning approaches, which often leads to blind and inefficient data augmentation. The authors propose a “programming with data” paradigm that treats structured knowledge representations as a unified foundation for both training and evaluation, drawing an analogy to software development: training data serve as source code, model training as compilation, evaluation as unit testing, and data refinement as debugging. This framework enables precise, concept- and reasoning-chain–oriented model repair through structured knowledge extraction, test-driven data engineering, concept-level gap analysis, and diagnosis of broken reasoning chains. Validated across 16 disciplines, the approach significantly enhances model performance without compromising general capabilities, and the authors release an open-source knowledge base, evaluation suite, and training corpora to support reproducibility and further research.
This study addresses the lack of a systematic review on the application of large language models (LLMs) in software engineering documentation and modeling tasks. Through a comprehensive literature survey, it establishes a multi-dimensional taxonomy that categorizes existing research by task type, offering an in-depth analysis of key technical approaches—including prompt engineering, natural language understanding, and structured language processing. The work further synthesizes the distribution of tasks, evaluation metrics, human assessment methodologies, and commonly used datasets across major conferences in the field. By systematically mapping the research landscape and identifying prevailing technical trends, this paper provides a thorough reference and strategic guidance for future investigations at the intersection of LLMs and software engineering.