taxonomy mapping

Designs, builds, and evaluates mapping functions and pipelines that convert labels from a source taxonomy to a target taxonomy and relabel datasets accordingly. Work includes creating rule-based or multi-stage relabeling pipelines, resolving conflicting or ambiguous mappings, implementing taxonomy-guided annotation workflows, and measuring automatic conversion rate, coverage, and correctness.

taxonomymapping

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of current expert-validated “LLM+script” workflows, which lack adaptability, cannot dynamically evolve based on feedback, and offer no effective pathway toward agent-based architectures. To overcome these challenges, the paper proposes a reversible “Strangler Fig” migration framework that transforms static workflows into composable, typed, and auditable stages. It introduces a three-tier convertibility classification—A/B/C—to enable dynamic routing and progressive evolution. This approach uniquely facilitates a smooth, structured transition from legacy LLM workflows to self-evolving agent systems while providing an assessment capability to determine evolutionary readiness.

adaptabilitylegacy systemsLLM workflows

This work addresses the challenge of efficiently constructing a comprehensive and well-structured taxonomy of artificial intelligence skills and tasks from massive hiring data. To this end, the authors propose TaxonomyBuilder, a framework that integrates systematic data filtering, clustering algorithms, and large language model–enhanced hierarchical label generation to automatically derive domain-specific taxonomies from curated, high-quality data subsets. Experimental results demonstrate that taxonomies built from filtered data exhibit significantly broader coverage and superior structural coherence compared to those generated from raw, unfiltered data using existing methods. The study thus establishes a novel paradigm for data-driven, automated taxonomy construction in specialized domains.

AI skillsdata filteringhierarchical taxonomy

TaxoAlign: Scholarly Taxonomy Generation Using Language Models

Oct 20, 2025
AL
Avishek Lahiri
🏛️ Indian Association for the Cultivation of Science | IT:U Interdisciplinary Transformation University Austria

Existing automated academic survey generation methods lack systematic structural alignment between generated taxonomies and those crafted by human experts, resulting in insufficient semantic coherence and hierarchical consistency. To address this, we propose TaxoAlign—a three-stage taxonomy generation framework jointly guided by topic modeling and instruction tuning, integrating topic-aware representation learning, large language model–driven hierarchical expansion, and structure-aware refinement. To enable rigorous evaluation, we introduce CS-TaxoBench, the first high-quality, expert-annotated benchmark for computer science taxonomies, and design the first automated, quantitative evaluation framework for structural alignment and semantic coherence. Experimental results demonstrate that TaxoAlign significantly outperforms all baselines in both automated metrics and human evaluations, achieving breakthrough improvements in hierarchical plausibility, cross-level semantic consistency, and expert alignment.

Automating scholarly taxonomy generation using language modelsBridging the gap between human and machine-created taxonomiesEvaluating structural alignment and semantic coherence of taxonomies

In domain ontology construction, mapping multi-source terminology to foundational concepts faces three key challenges: high cost and subjectivity of manual approaches, shallow semantic modeling and poor cross-domain consistency of automated methods, and weak interpretability. To address these, this paper proposes the first LLM-driven framework integrating expert calibration with iterative prompt optimization. The framework combines expert-guided annotation, multi-stage prompt engineering, and a human-in-the-loop validation cycle to generate concept links with high confidence and full interpretability. Evaluated on the concept necessity mapping task, it achieves an F1-score of 0.97—substantially surpassing the human baseline (0.68)—and marks the first instance of scalable ontology alignment that simultaneously attains expert-level accuracy and transparent, auditable reasoning.

Automating taxonomy alignment to replace costly manual expert reviewEnhancing transparency and consistency in automated concept mapping processesImproving semantic relationship handling in cross-domain taxonomy mapping

CodeTaxo: Enhancing Taxonomy Expansion with Limited Examples via Code Language Prompts

Aug 17, 2024
QZ
Qingkai Zeng
🏛️ University of Notre Dame | University of Washington

Identifying appropriate parent classes for novel concepts in small-scale existing taxonomies (<100 nodes) remains challenging due to sparse structural signals and lack of labeled training data. Method: This paper proposes a label-free, large language model (LLM)-driven taxonomy expansion method. Its core innovation is *code-style prompting*: explicitly encoding hierarchical semantic structure via indentation, nesting, and functional abstraction—programming conventions that enable zero- or few-shot LLM comprehension and relational reasoning over taxonomy hierarchies. The approach integrates taxonomy-aware input representations with structured code prompts, eliminating reliance on large-scale annotated or self-supervised data construction. Results: Evaluated on five cross-domain real-world benchmarks, the method achieves an average 12.6% absolute improvement in parent-class prediction accuracy over state-of-the-art methods. Gains are especially pronounced in extremely small-scale settings (30–80 nodes), demonstrating robustness where conventional supervised or embedding-based approaches falter.

Expanding taxonomies with limited existing examplesImproving accuracy for small taxonomies under 100 entitiesLeveraging code prompts to enhance structural understanding

Latest Papers

What's happening recently
View more

This study addresses the high cost and expert dependency of manual taxonomy construction in software engineering (SE) by conducting the first systematic, multi-dimensional empirical evaluation of large language model (LLM)-driven automatic classification in this domain. Leveraging two representative approaches—TnT-LLM and CLIMB—and five state-of-the-art LLMs across seven human-annotated SE paper datasets, the work analyzes performance along key dimensions including classification quality, alignment with expert judgments, reliability, and efficiency. Results reveal that TnT-LLM achieves near-human classification quality but incurs high computational cost and structural complexity, whereas CLIMB offers 15–40× faster inference and 8–49× lower cost at the expense of reduced accuracy in tasks requiring deep technical reasoning. The findings elucidate critical trade-offs among quality, cost, and complexity, providing actionable guidance for method selection in practice.

automated methodsempirical evaluationLarge Language Models

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study addresses the challenges of maintaining consistency across heterogeneous schema languages—such as JSON Schema, XSD, and SHACL—during multilingual data model evolution, where fragmented converters, variable quality, and information loss impede reliable interoperability. The work proposes a novel approach that models schema languages and black-box converters as nodes and directed edges in a graph, enabling composable and evaluable conversion path orchestration. By integrating graph-based search, quality-aware ranking (combining agent-assisted and human evaluation), and failure backtracking, the method supports automated, reproducible cross-language schema transformation. The resulting open-source toolchain, Schema Conversion Orchestrator, integrated into the MetaConfigurator platform, successfully produced valid outputs for 43 out of 60 real-world tasks and precisely identified missing ecosystem components in the remaining 17, thereby delineating the current boundaries of schema conversion capabilities.

black-box convertersconverter orchestrationdata model consistency

Hot Scholars

SV

Svitlana Volkova

Chief of AI, Office of Science and Technology, Aptima Inc.
Artificial IntelligenceMachine LearningComputational Social Science
MY

Minglai Yang

CS Undergraduate student, University of Arizona
Natural Language ProcessingLarge Language ModelsMachine Learning
XG

Xinyu Guo

Samsung Research America
AIcomputer visionmachine learningmedical image analysis