information extraction

Designs, implements, and evaluates systems and pipelines that identify and normalize entities, extract relations and facts, and convert unstructured or semi-structured text and tables into structured records or knowledge-graph triples. This includes building and integrating models, parsers, labelers and extraction libraries for document- and table-level extraction, entity and relation extraction models, and end-to-end information/knowledge extraction pipelines.

informationextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$186K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

TEXT2DB: Integration-Aware Information Extraction with Large Language Model Agents

Oct 27, 2025
YJ
Yizhu Jiao
🏛️ University of Illinois Urbana-Champaign

Information extraction (IE) outputs often mismatch downstream database schemas, hindering direct integration. Method: This paper introduces TEXT2DB—a novel task requiring models to dynamically perform data completion, row insertion, and column expansion based on user instructions, document collections, and target database schemas. To address it, we propose OPAL, an agent framework operating via an Observe-Plan-Analyze closed-loop that orchestrates database interaction, code generation, IE model invocation, and pre-execution feedback analysis for end-to-end instruction understanding, schema alignment, and structured data population. Contribution/Results: Experiments demonstrate that OPAL accurately executes complex IE-database joint tasks across diverse database systems, significantly improving information-to-database deployment efficiency. The study further identifies critical challenges—including large-scale schema adaptation and model hallucination—highlighting open research directions for robust database-grounded IE.

Aligning information extraction with target database schema requirementsExtracting structured knowledge from text for database integrationHandling user instructions for dynamic database updates from documents

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Oct 28, 2024
QZ
Qintong Zhang
🏛️ Shanghai Artificial Intelligence Laboratory | Peking University

This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.

Address challenges in layout detection and multi-modal data integrationConvert unstructured documents into structured machine-readable dataImprove parsing accuracy for complex layouts and high-density text

This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.

data curationentity identityentity resolution

LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows

Oct 06, 2025
SA
Samy Ateia
🏛️ University of Regensburg | University of Bayreuth

With the exponential growth of scientific literature, automated extraction of key concepts remains challenging, particularly due to poor cross-disciplinary adaptability. Method: This paper proposes a lightweight LLM-based semantic extraction method supporting FAIR implementation in scholarly workflows. It introduces a context learning–driven zero-/few-shot domain adaptation mechanism that enables rapid, fine-tuning–free adaptation to new disciplines. We systematically benchmark multiple open-source and commercial LLMs on concept identification tasks and develop an interactive online prototype system. Contribution/Results: Empirical evaluation in computer science—complemented by user studies—demonstrates the method’s effectiveness in structured literature review, knowledge graph construction, and information retrieval. It significantly improves both accuracy and cross-domain generalization of concept extraction, offering a scalable technical pathway for intelligent, full-lifecycle scholarly knowledge services.

Enabling rapid domain adaptation for scientific information extractionExtracting key concepts from scientific documents using LLMsSupporting FAIR principles in scientific publishing workflows

Automated Requirements Relation Extraction

Jan 22, 2024
QM
Quim Motger
🏛️ Universitat Polit`ecnica de Catalunya

To address challenges in requirements engineering—including difficulty in identifying relationships among natural language requirements, high manual annotation costs, and poor domain adaptability—this paper proposes an NLP-driven, systematic relation extraction framework. It is the first to integrate a requirements relationship ontology with multi-paradigm NLP techniques: dependency parsing, semantic role labeling, named entity recognition, BERT-based supervised fine-tuning, and retrieval-augmented methods. A unified classification-based evaluation framework is established to clarify core challenges and evolutionary pathways. The framework supports major requirement relations (e.g., *refines*, *conflicts*) and enables reusable, extensible relation modeling. Experimental results demonstrate significant improvements in automation capability and accuracy for large-scale adaptive requirements management systems, thereby strengthening requirements evolution analysis and consistency verification.

Addressing ambiguity and effort in requirements engineeringAutomated extraction of relations between textual requirementsExploring NLP techniques for efficient relation identification

Latest Papers

What's happening recently
View more

This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.

data integrationknowledge representationlarge language models

This study addresses the challenge of effectively integrating structured data with unstructured text, a longstanding barrier in data management. It presents the first systematic argument for the necessity of textual data integration and introduces a unified framework that synergistically combines natural language processing, knowledge extraction, and traditional data integration techniques. By leveraging semantic alignment, the framework achieves deep integration between textual content and structured schemas, thereby tackling key challenges inherent in heterogeneous data integration. The work comprehensively surveys existing methodologies and outstanding issues, establishing a theoretical foundation for the emerging field of textual data integration and offering clear guidance for future research and practical implementation.

Data IntegrationHeterogeneous DataStructured Data

This work proposes DySECT, the first dynamic information extraction system that enables co-evolution of knowledge and extraction to address challenges such as the dynamic evolution of domain-specific terminology, lagging expert taxonomies, and difficulties in recognizing rare terms. DySECT continuously constructs a knowledge base by extracting triples using large language models, integrates probabilistic knowledge representation with graph-based reasoning to support autonomous knowledge expansion, and enhances the extraction model through prompt tuning, few-shot learning, or fine-tuning on synthetic data—establishing a closed-loop “extraction–knowledge” reinforcement mechanism. Experiments in dynamic domains including healthcare, legal, and human resources demonstrate that DySECT significantly improves both the accuracy and timeliness of information extraction, achieving continuous self-optimization of system capabilities.

domain-specific terminologydynamic adaptationinformation extraction

This work addresses the limitations of traditional knowledge graph construction approaches, wherein structural decisions are hard-coded into rigid pipelines, resulting in tight coupling between schema and construction process and hindering support for ontology-level tasks. To overcome this, the authors propose an ontology-oriented construction framework featuring a novel intrinsic-relational routing mechanism. This mechanism dynamically assigns attributes to corresponding schema modules through iterative attribute classification, enabling a declarative and reusable decoupled design. The pipeline integrates rule-based cleaning, tool-augmented large language model–assisted annotation, and human review. Evaluated on Wikidata (January 2026), the resulting graph comprises 34 million nodes and 61.2 million edges, achieving 93.3% schema coverage and 98.0% module assignment accuracy, effectively supporting five ontology-level applications.

knowledge graphontologyproperty graph

This work addresses the challenge of automatically constructing knowledge graphs from multi-source, heterogeneous, and unstructured textual data by proposing an interpretable and interoperable approach that integrates generative AI with semantic web technologies. The method supports adaptive alignment across diverse text types and schema specifications, leveraging natural language processing, information extraction, and causal modeling to enable end-to-end knowledge graph construction. Domain-specific knowledge graphs were developed and validated in three real-world scenarios—news and social media, architectural engineering operations documentation, and electronic health records—yielding tailored algorithms, benchmark evaluations, and in-depth analytical insights. These resources effectively support applications such as digital transformation discourse analysis, scientific trend identification, and causal reasoning in biomedical research.

InteroperabilityKnowledge Graph ConstructionScalable Information Extraction

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
JH

Jiawei Han

Abel Bliss Professor of Computer Science, University of Illinois
data miningdatabase systemsdata warehousinginformation networks
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
WG

Wenbo Guo

UC Santa Barbara
Machine LearningSecurity
HQ

Huamin Qu

Chair Professor, Hong Kong University of Science and Technology
Data visualizationHuman-Computer InteractionExplainable AIE-Learning