cross-domain corpus alignment

Design and build methods, algorithms, and evaluation processes that align and associate concepts, records, and taxonomic nodes across heterogeneous corpora and disparate domains, producing correspondence models and transformation procedures that reconcile differing schemas, terminologies, and granularities. This includes resolution-aware mapping (node-to-node, node-to-cluster), cross-domain association measures, and tools to enable cross-corpus retrieval and data synthesis.

cross-domaincorpusalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Words as Bridges: Exploring Computational Support for Cross-Disciplinary Translation Work

Mar 24, 2025
CB
Calvin Bao
🏛️ University of Maryland | MIT

Interdisciplinary literature exploration is often impeded by terminological barriers across domains. This paper conceptualizes disciplines as heterogeneous linguistic communities and, for the first time, adapts unsupervised cross-lingual word embedding alignment techniques to interdisciplinary concept alignment—preserving domain-specific terms as cognitive bridges rather than eliminating or oversimplifying them. Our method comprises: (1) domain-specific word vector training; (2) unsupervised alignment of embedding spaces across disciplines; and (3) construction and interactive design of a prototype concept-level cross-domain search engine. Evaluated in two case studies, our approach demonstrates effective semantic mapping, enabling concept-level—rather than lexical-level—cross-domain retrieval. Key contributions are: (1) establishing a novel paradigm for interdisciplinary concept alignment; (2) empirically validating domain terms as computationally tractable conceptual anchors; and (3) providing evidence-based insights for scholar-centered information-seeking interface design.

Aligning domain-specific word embeddings for cross-disciplinary conceptual explorationBridging jargon gaps between disciplines for better translationDeveloping computational tools to aid cross-domain information seeking

Existing corpus construction methods produce only flat document collections, lacking systematic knowledge organization and thus failing to meet the demand of large language models for high-quality structured data. This work proposes the CORTEX framework, which introduces the first three-layer heterogeneous Ontology-based Corpus Graph (OCG), comprising a quality-optimized content layer, a lightweight ontology layer evolved via LLM-driven mechanisms, and a cross-domain alignment layer. This architecture enables cross-domain associations at arbitrary granularities and supports automatic ontology evolution. Leveraging this framework, we construct a structured, high-quality corpus of 24.14 billion tokens and release CortexBench, a benchmark for cross-domain retrieval and reasoning, demonstrating its effectiveness across eight state-of-the-art large language models.

corpus organizationcross-domain alignmentknowledge structure

Setting The Table with Intent: Intent-aware Schema Generation and Editing for Literature Review Tables

Jul 18, 2025
VP
Vishakh Padmakumar
🏛️ New York University | AI2 | Northwestern University

The exponential growth of academic literature poses significant challenges for efficiently constructing comparative tables in survey papers. Existing schema generation methods suffer from ambiguous evaluation criteria and limited editability. To address these issues, this paper proposes an intent-aware schema generation and editing framework: (1) it introduces intent modeling to mitigate semantic ambiguity in comparative dimension identification; (2) it designs an editable generation pipeline enabling on-demand customization of comparison dimensions; (3) it constructs the first benchmark dataset tailored for conditional schema generation; and (4) it integrates LLM-based prompt engineering with lightweight fine-tuning, combining one-shot generation and multi-stage editing strategies. Experimental results demonstrate that intent enhancement substantially improves schema reconstruction accuracy, while the editing mechanism further refines output quality. Notably, our lightweight fine-tuned model achieves performance competitive with state-of-the-art prompting-based large language models.

Addressing ambiguity in schema evaluation through synthesized intentsDeveloping refinement methods to improve generated schemasGenerating schemas to organize academic literature collections

Magneto: Combining Small and Large Language Models for Schema Matching

Dec 11, 2024
YL
Yurong Liu
🏛️ New York University

Schema matching across heterogeneous, cross-source data faces dual bottlenecks: small language models (SLMs) rely heavily on scarce labeled data, while large language models (LLMs) incur prohibitive computational costs and suffer from context-length limitations. Method: We propose a two-stage SLM–LLM collaborative framework: SLMs perform efficient candidate generation, followed by LLM-based re-ranking via prompt engineering and generative self-supervised fine-tuning. Contribution/Results: This work introduces the first SLM–LLM collaboration paradigm for schema matching; designs a novel generative self-supervised fine-tuning strategy for LLMs—eliminating dependence on manual annotations; and establishes BioSchema, the first realistic, biomedical-domain-specific schema matching benchmark. Evaluated across multiple domains, our method achieves state-of-the-art accuracy while reducing inference cost by 42%, significantly outperforming both pure-SLM and pure-LLM baselines.

Combining SLMs and LLMs for efficient schema matchingGenerating synthetic training data for self-supervised SLM fine-tuningReducing computational costs without losing accuracy

Scientific document retrieval faces significant challenges due to the scarcity of domain-specific labeled data and the highly specialized nature of technical terminology, which often leads existing methods to suffer from conceptual redundancy or insufficient coverage. To address these limitations, this work proposes an academic concept indexing framework that integrates a structured scholarly taxonomy with large language models to extract and organize key concepts. The framework introduces two novel mechanisms: Concept-Coverage-aware Query Generation (CCQGen) and Concept-Focused Context Expansion (CCExpand), which jointly enhance the retrieval system’s capacity to understand and match scientific semantics. Experimental results demonstrate that the proposed approach substantially improves query quality, concept alignment, and overall retrieval effectiveness, outperforming current state-of-the-art methods on scientific document retrieval benchmarks.

academic conceptcontext augmentationdomain adaptation

Latest Papers

What's happening recently
View more

This study addresses the challenge of aligning and interpreting knowledge structures across heterogeneous textual corpora by proposing a term-centric, hierarchical knowledge construction framework. Departing from conventional full-document representations, the approach maps multi-source documents into a shared semantic space through automated term extraction and integrates domain priors with data-driven clustering to generate interpretable knowledge hierarchies. Evaluated on a newly introduced benchmark comprising over one million English–German multi-source documents, the method significantly improves cross-source consistency and hierarchy quality. Its practical utility is further demonstrated through the successful construction of regional innovation technology maps for Germany, highlighting its applicability in policy analysis and innovation monitoring.

cross-source alignmentheterogeneous corporaknowledge organization

Existing approaches to scientific knowledge classification often suffer from semantic inconsistency and structural misalignment within hierarchical taxonomies, hindering their ability to effectively organize the rapidly expanding body of scholarly literature. This work proposes a hierarchical classification framework grounded in large language models, which integrates a bidirectional title generation mechanism—combining bottom-up abstraction with top-down constraints—to jointly optimize vertical alignment across levels and horizontal semantic coherence among sibling nodes. The method explicitly models semantic dependencies among nodes at the same hierarchy level, significantly enhancing the logical structure, semantic fidelity, and title quality of the resulting taxonomy. Extensive experiments demonstrate strong performance across multiple benchmark datasets and reveal robust cross-lingual generalization capabilities, particularly on Chinese scientific literature.

hierarchical structurelarge language modelsscientific literature

This work addresses the inconsistency of cross-document software mentions in scientific literature by proposing a normalization framework that integrates semantic embeddings, knowledge base retrieval, and density-based clustering. The approach combines Sentence-BERT for semantic representation with FAISS for efficient retrieval, and incorporates surface form normalization and abbreviation resolution to handle out-of-vocabulary and ambiguous mentions. To enhance scalability in large-scale settings, an entity-type-aware blocking strategy is introduced. Guided by semantic centroids, HDBSCAN clustering enables the system to achieve CoNLL F1 scores of 0.98, 0.98, and 0.96 on the three subtasks of the SOMD 2026 shared task, substantially outperforming baseline methods.

Cross-Document Coreference ResolutionEntity LinkingScientific Corpora

Improving LLM-based Ontology Matching with fine-tuning on synthetic data

Nov 27, 2025
GS
Guilherme Sousa
🏛️ IRIT | Université de Toulouse 2 Jean Jaurès | Universidade Federal Rural de Recife | Univ. Grenoble Alpes | Inria | CNRS | Grenoble INP

Ontology matching in zero-shot settings faces challenges of semantic gap and combinatorial explosion in the search space. This paper proposes a module-level, LLM-based direct alignment method: first, domain-informed heuristic rules prune the candidate module-pair search space; second, an automated prompting mechanism guides large language models to generate high-quality, diverse synthetic alignment corpora; finally, the LLM is lightweight fine-tuned on this corpus. The core innovation lies in the tight coupling of search-space pruning and LLM-driven synthetic data generation, forming a closed-loop fine-tuning paradigm. Experiments across multiple benchmark datasets from the OAEI complex track demonstrate that the fine-tuned model significantly outperforms zero-shot baselines, achieving an average 12.7% improvement in Top-1 accuracy—validating both effectiveness and generalizability.

Enhancing LLM performance in ontology matching via fine-tuningGenerating synthetic datasets to address training data scarcityReducing search space for efficient ontology alignment generation

CMOMgen: Complex Multi-Ontology Alignment via Pattern-Guided In-Context Learning

Oct 24, 2025
MC
Marta Contreiras Silva
🏛️ Universidade de Lisboa | INESC-ID | Instituto Superior Técnico

This work addresses the coarse-grained semantic integration problem in multi-ontology settings by introducing the Complex Multi-Ontology Matching (CMOM) task: mapping each source entity to a logical expression—e.g., conjunction or disjunction—over multiple target entities, enabling fine-grained equivalence modeling and traceable, hierarchical alignment. To this end, we propose CMOMgen, the first end-to-end framework supporting arbitrary numbers of target ontologies and entities. CMOMgen innovatively integrates schema-guided in-context learning with retrieval-augmented generation (RAG): it retrieves semantically relevant ontology classes and reference alignments via analogical retrieval, guiding large language models to generate semantically sound and logically consistent composite mappings. Evaluated on three biomedical datasets, CMOMgen achieves F1 scores of 63%–79%, establishing new state-of-the-art performance on two benchmark tasks. Human evaluation further confirms its generalizability and practicality: 46% of non-reference mappings received the highest rating.

Aligns source entities to composite logical expressionsEnables multi-ontology matching without target restrictionsGenerates semantically sound mappings for knowledge graphs

Hot Scholars

YY

Yun Ye

Intel
Computer VisionDeep LearningSemiconductor Physics
XS

Xuesong Shi

Galbot
robotic visionheterogeneous computinggraph signal processingSLAM
JW

Jingya Wang

Assistant Professor, ShanghaiTech University
Computer VisionEmbodied AIHuman-Object Interaction