llm fact extraction

Designs and implements pipelines, prompt templates, and validation tooling that use large language models to extract and structure factual knowledge or priors from unstructured inputs, producing typed and labeled facts. Work covers per-unit processing (for example, per-function), deriving dataflow facts, variable types and invariants, and serializing or formatting those facts for logical ingestion and downstream analysis (e.g., datalog), plus methods to curate and validate the extracted knowledge.

llmfactextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Lost in the Pipeline: How Well Do Large Language Models Handle Data Preparation?

Nov 17, 2025
MS
Matteo Spreafico
🏛️ Politecnico di Milano

Existing automated data preparation tools lack robust semantic understanding and struggle with complex, context-dependent data quality issues. Method: This study investigates the efficacy of large language models (LLMs) in data profiling and cleaning on low-quality datasets. We propose a customized data quality assessment framework informed by a practitioner-focused user study, and systematically evaluate both general-purpose and fine-tuned table-centric LLMs—via prompt engineering—on tasks including anomaly detection, cleaning logic generation, and error repair, benchmarking against traditional tools (e.g., Trifacta, OpenRefine). Contribution/Results: LLMs significantly outperform conventional tools in contextual reasoning and generating interpretable, human-verifiable cleaning rules; however, their output precision and deterministic verifiability remain limited. This work establishes the first evaluation paradigm specifically designed for LLMs in data preparation and empirically validates their viability—and practical boundaries—as collaborative “data engineering partners.”

Assessing LLMs' ability in data profiling and cleaningComparing LLM support with traditional data preparation toolsEvaluating LLMs' effectiveness in data preparation tasks

This work addresses the challenges of accuracy and scalability in knowledge graph fact verification at scale, where existing automated methods remain immature. We propose the first multidimensional evaluation framework for assessing large language models (LLMs) in this context, systematically examining their capabilities along three dimensions: internal knowledge, retrieval-augmented generation (RAG), and multi-model consensus. To support this evaluation, we construct a RAG dataset comprising two million documents, a FactCheck benchmark, and an interactive analysis platform, conducting experiments across three real-world knowledge graphs. Our results reveal that while LLMs show promise, their performance lacks robustness; furthermore, the effectiveness of RAG and multi-model strategies varies significantly, underscoring both the necessity of systematic evaluation and the practical utility of our proposed framework.

BenchmarkingFact VerificationFactual Accuracy

From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs

May 29, 2025
XG
Xuan Gong
🏛️ Tongji University | Shanghai Jiao Tong University

This paper addresses the factual inconsistency gap in fine-tuned large language models (LLMs) between known (in-distribution) and unknown (out-of-distribution) knowledge. Methodologically, it establishes— for the first time at the theoretical level—that test-time prompting techniques (e.g., in-context learning [ICL] and chain-of-thought [CoT]) can attenuate or even override the influence of fine-tuning data; it further introduces the “prompt dominance” theory, redefining evaluation criteria for fine-tuning data. Empirical results demonstrate that ICL and CoT improve factual accuracy by over 35% on out-of-distribution knowledge tasks. The core contribution lies in uncovering the compensatory mechanism whereby prompt engineering mitigates fine-tuning biases, thereby establishing a novel paradigm—“prompting to compensate for fine-tuning deficiencies.” This work provides interpretable, quantifiable theoretical principles and practical guidelines for synergistically optimizing fine-tuning and reasoning.

Exploring interaction between fine-tuning data and promptsMitigating factuality gap via inference-stage strategiesUnderstanding the factuality gap in fine-tuned LLMs

Tasks People Prompt: A Taxonomy of LLM Downstream Tasks in Software Verification and Falsification Approaches

Apr 14, 2024
VB
V. Braberman
🏛️ Universidad de Buenos Aires | CONICET-Universidad de Buenos Aires | The University of Manchester | Universidade Federal do Amazonas

Current LLM-native software engineering lacks a systematic practical framework—particularly in verification and falsification—necessitating unified task taxonomies and prompt-pattern conceptualizations. Method: We conduct a systematic literature review of over 100 papers, employing bibliometric analysis and conceptual clustering to map, classify, and abstract LLM-based downstream tasks in software engineering (SE). Contribution/Results: We propose the first fine-grained SE-specific taxonomy for LLM downstream tasks, encompassing six core clusters: testing, fuzzing, bug localization, vulnerability detection, static analysis, and program verification. Our taxonomy uniquely balances cross-task abstraction with task-specific variation modeling, uncovering generalizable prompt-engineering principles. It provides a foundational framework for targeted LLM adaptation, benchmark construction, and empirically grounded engineering practice in SE.

Developing conceptual frameworks for LLM-native software engineering practicesIdentifying compositional patterns for reliable LLM-native system designSystematically analyzing generative transformations in software verification

Large language models (LLMs) frequently fail in real-world tool invocation due to intent misinterpretation, incorrect parsing of tool documentation, and parameterization errors. To address this, we propose a curriculum-inspired structured reasoning framework that replaces free-form chain-of-thought prompting with guided, template-based reasoning—explicitly decoupling the process into three sequential stages: *intent parsing*, *tool matching*, and *parameter generation*. Our framework employs stepwise structured prompts to jointly model user goals and tool functionalities, thereby enhancing invocation robustness and decision interpretability. Evaluated across multiple state-of-the-art models (e.g., LLaMA-3, Qwen2) and benchmarks (ToolBench, API-Bank), it reduces relative error rates by 3–12% over strong baselines. The core contribution lies in transforming implicit, unstructured reasoning into an explicit, traceable, and modular pipeline—balancing accuracy with transparency and auditability.

Addresses incorrect parameterization and poor tool selection in LLMsImproves incomplete understanding of user goals and tool documentationSolves misinterpretation of user intent in function-calling tasks

Latest Papers

What's happening recently
View more

This work addresses the challenges of scarce labeled data and weakly expressed, low-salience product attributes in applications such as digital product passports by proposing a two-step verification generative information extraction framework that integrates pretrained language models (PLMs) with large language models (LLMs). The approach first employs a PLM for initial candidate extraction and then leverages a locally deployable open-source LLM—such as those in the Llama family—for secondary verification and error correction, substantially improving extraction accuracy for sparse and weakly expressed entities. Experimental results demonstrate that the proposed framework enhances generalization capability and enables medium-scale models to approach the performance of much larger models, all while preserving data privacy and maintaining computational efficiency. The method has been successfully integrated into a demonstration system tailored for digital product passports.

digital product passportinformation extractionlarge language models

This study addresses the lack of systematic evaluation for Datalog programs generated by large language models (LLMs) by constructing a benchmark comprising 136 tasks and proposing reliable metrics based on execution verification and mutation analysis. Through experiments involving six prompting strategies applied to six LLMs and coding agents, results indicate that direct prompting achieves a maximum match rate of 68.4%, whereas coding agents attain 83.8% while effectively eliminating most compilation errors. The research precisely identifies semantic and compilation errors, revealing that recursive reasoning remains a core challenge for current models. Overall, this work establishes a new paradigm for evaluating the logical programming capabilities of LLMs.

Benchmark evaluationDatalogLarge Language Models

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

Dec 18, 2025
HL
Hao Liang
🏛️ Peking University | Institute for Advanced Algorithms Research | OriginHub Technology | OpenDataLab | Shanghai Artificial Intelligence Laboratory | LLaMA-Factory Team

The era of large language models (LLMs) faces critical challenges including insufficient high-quality data supply, fragmented data preparation pipelines, poor reproducibility, and lack of model-in-the-loop support. Method: We propose the first LLM-driven, unified data preparation framework for data-centric AI, featuring system-level abstractions and PyTorch-style APIs for modular design. We introduce DataFlow-Agent—the first agent that synthesizes executable data pipelines end-to-end from natural language specifications—and integrate LLM-powered operator synthesis, iterative validation, 200+ reusable operators, and six domain-agnostic pipeline templates. Results: Experiments on Text-to-SQL, code generation, and mathematical reasoning show our synthesized data significantly outperforms human-annotated and domain-specific synthetic data. Remarkably, just 10K samples surpass the performance of models trained on the million-scale Infinity-Instruct dataset, empirically validating the decisive impact of data quality on model performance.

Addresses scalable, reliable data preparation for LLMsAutomates pipeline creation from natural language specificationsReplaces ad-hoc scripts with modular, reusable data transformations

This work addresses the lack of traceability in existing domain-specific fine-tuning approaches, which often leads to blind and inefficient data augmentation. The authors propose a “programming with data” paradigm that treats structured knowledge representations as a unified foundation for both training and evaluation, drawing an analogy to software development: training data serve as source code, model training as compilation, evaluation as unit testing, and data refinement as debugging. This framework enables precise, concept- and reasoning-chain–oriented model repair through structured knowledge extraction, test-driven data engineering, concept-level gap analysis, and diagnosis of broken reasoning chains. Validated across 16 disciplines, the approach significantly enhances model performance without compromising general capabilities, and the authors release an open-source knowledge base, evaluation suite, and training corpora to support reproducibility and further research.

data engineeringdomain specializationknowledge transfer

This study addresses the lack of a systematic review on the application of large language models (LLMs) in software engineering documentation and modeling tasks. Through a comprehensive literature survey, it establishes a multi-dimensional taxonomy that categorizes existing research by task type, offering an in-depth analysis of key technical approaches—including prompt engineering, natural language understanding, and structured language processing. The work further synthesizes the distribution of tasks, evaluation metrics, human assessment methodologies, and commonly used datasets across major conferences in the field. By systematically mapping the research landscape and identifying prevailing technical trends, this paper provides a thorough reference and strategic guidance for future investigations at the intersection of LLMs and software engineering.

Generative AILarge Language ModelsSoftware Documentation

Hot Scholars

XJ

Xiongnan Jin

Shenzhen University
Knowledge GraphLarge Language ModelsData MiningSemantic Computing
PL

Peiyang Liu

Peking University
Information RetrievalLarge Language Model
JJ

Jiwei Jiang

Huazhong University Of Science And Technology
Computer Vision
VN

Varun Nagaraj Rao

Center for Information Technology Policy, Princeton University
AI AuditsLaborHCIVision-Language Models
SZ

Shuxin Zheng

Deputy Director, Zhongguancun Institute of Artificial Intelligence
General AIGenerative AI