Score
Designs, implements, and evaluates systems and pipelines that identify and normalize entities, extract relations and facts, and convert unstructured or semi-structured text and tables into structured records or knowledge-graph triples. This includes building and integrating models, parsers, labelers and extraction libraries for document- and table-level extraction, entity and relation extraction models, and end-to-end information/knowledge extraction pipelines.
Text-to-structured generation (e.g., tables, knowledge graphs, charts) for agent-centric AI is a foundational infrastructure enabling context-aware retrieval and autonomous reasoning, yet suffers from fragmented methodologies, scarce standardized datasets, and inconsistent evaluation protocols. Method: We conduct a systematic literature review integrating techniques from NLP, information extraction, knowledge representation, and machine learning to establish the first holistic analytical framework—comprising task taxonomy, benchmark dataset inventory, and unified evaluation metrics. Contribution/Results: We introduce the first general-purpose evaluation framework for structured output generation, explicitly identifying methodological limitations and core challenges (e.g., fidelity, composability, and reasoning-aware assessment). We comprehensively map research gaps and affirm the centrality of this direction in next-generation AI systems, providing both theoretical grounding and practical guidance for future algorithmic development and empirical validation.
Information extraction (IE) outputs often mismatch downstream database schemas, hindering direct integration. Method: This paper introduces TEXT2DB—a novel task requiring models to dynamically perform data completion, row insertion, and column expansion based on user instructions, document collections, and target database schemas. To address it, we propose OPAL, an agent framework operating via an Observe-Plan-Analyze closed-loop that orchestrates database interaction, code generation, IE model invocation, and pre-execution feedback analysis for end-to-end instruction understanding, schema alignment, and structured data population. Contribution/Results: Experiments demonstrate that OPAL accurately executes complex IE-database joint tasks across diverse database systems, significantly improving information-to-database deployment efficiency. The study further identifies critical challenges—including large-scale schema adaptation and model hallucination—highlighting open research directions for robust database-grounded IE.
This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.
This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.
With the exponential growth of scientific literature, automated extraction of key concepts remains challenging, particularly due to poor cross-disciplinary adaptability. Method: This paper proposes a lightweight LLM-based semantic extraction method supporting FAIR implementation in scholarly workflows. It introduces a context learning–driven zero-/few-shot domain adaptation mechanism that enables rapid, fine-tuning–free adaptation to new disciplines. We systematically benchmark multiple open-source and commercial LLMs on concept identification tasks and develop an interactive online prototype system. Contribution/Results: Empirical evaluation in computer science—complemented by user studies—demonstrates the method’s effectiveness in structured literature review, knowledge graph construction, and information retrieval. It significantly improves both accuracy and cross-domain generalization of concept extraction, offering a scalable technical pathway for intelligent, full-lifecycle scholarly knowledge services.
To address challenges in requirements engineering—including difficulty in identifying relationships among natural language requirements, high manual annotation costs, and poor domain adaptability—this paper proposes an NLP-driven, systematic relation extraction framework. It is the first to integrate a requirements relationship ontology with multi-paradigm NLP techniques: dependency parsing, semantic role labeling, named entity recognition, BERT-based supervised fine-tuning, and retrieval-augmented methods. A unified classification-based evaluation framework is established to clarify core challenges and evolutionary pathways. The framework supports major requirement relations (e.g., *refines*, *conflicts*) and enables reusable, extensible relation modeling. Experimental results demonstrate significant improvements in automation capability and accuracy for large-scale adaptive requirements management systems, thereby strengthening requirements evolution analysis and consistency verification.
This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.
This study addresses the challenge of effectively integrating structured data with unstructured text, a longstanding barrier in data management. It presents the first systematic argument for the necessity of textual data integration and introduces a unified framework that synergistically combines natural language processing, knowledge extraction, and traditional data integration techniques. By leveraging semantic alignment, the framework achieves deep integration between textual content and structured schemas, thereby tackling key challenges inherent in heterogeneous data integration. The work comprehensively surveys existing methodologies and outstanding issues, establishing a theoretical foundation for the emerging field of textual data integration and offering clear guidance for future research and practical implementation.
This work proposes DySECT, the first dynamic information extraction system that enables co-evolution of knowledge and extraction to address challenges such as the dynamic evolution of domain-specific terminology, lagging expert taxonomies, and difficulties in recognizing rare terms. DySECT continuously constructs a knowledge base by extracting triples using large language models, integrates probabilistic knowledge representation with graph-based reasoning to support autonomous knowledge expansion, and enhances the extraction model through prompt tuning, few-shot learning, or fine-tuning on synthetic data—establishing a closed-loop “extraction–knowledge” reinforcement mechanism. Experiments in dynamic domains including healthcare, legal, and human resources demonstrate that DySECT significantly improves both the accuracy and timeliness of information extraction, achieving continuous self-optimization of system capabilities.
This work addresses the limitations of traditional knowledge graph construction approaches, wherein structural decisions are hard-coded into rigid pipelines, resulting in tight coupling between schema and construction process and hindering support for ontology-level tasks. To overcome this, the authors propose an ontology-oriented construction framework featuring a novel intrinsic-relational routing mechanism. This mechanism dynamically assigns attributes to corresponding schema modules through iterative attribute classification, enabling a declarative and reusable decoupled design. The pipeline integrates rule-based cleaning, tool-augmented large language model–assisted annotation, and human review. Evaluated on Wikidata (January 2026), the resulting graph comprises 34 million nodes and 61.2 million edges, achieving 93.3% schema coverage and 98.0% module assignment accuracy, effectively supporting five ontology-level applications.
This work addresses the challenge of automatically constructing knowledge graphs from multi-source, heterogeneous, and unstructured textual data by proposing an interpretable and interoperable approach that integrates generative AI with semantic web technologies. The method supports adaptive alignment across diverse text types and schema specifications, leveraging natural language processing, information extraction, and causal modeling to enable end-to-end knowledge graph construction. Domain-specific knowledge graphs were developed and validated in three real-world scenarios—news and social media, architectural engineering operations documentation, and electronic health records—yielding tailored algorithms, benchmark evaluations, and in-depth analytical insights. These resources effectively support applications such as digital transformation discourse analysis, scientific trend identification, and causal reasoning in biomedical research.