taxonomy design

Design and build structured classification systems and controlled vocabularies—hierarchical or flat—that define coherent categories and relations for tasks, datasets, labels, attributes, interactions, systems, evaluations, conceptual constructs, and sources-of-authority. Specify category definitions, boundaries, metadata and mapping rules, produce annotation guidelines and item-to-category mappings, and use the taxonomy to enable organized analysis, benchmarking, gap identification, and unified cross-study comparisons.

taxonomydesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$169K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of efficiently constructing a comprehensive and well-structured taxonomy of artificial intelligence skills and tasks from massive hiring data. To this end, the authors propose TaxonomyBuilder, a framework that integrates systematic data filtering, clustering algorithms, and large language model–enhanced hierarchical label generation to automatically derive domain-specific taxonomies from curated, high-quality data subsets. Experimental results demonstrate that taxonomies built from filtered data exhibit significantly broader coverage and superior structural coherence compared to those generated from raw, unfiltered data using existing methods. The study thus establishes a novel paradigm for data-driven, automated taxonomy construction in specialized domains.

AI skillsdata filteringhierarchical taxonomy

10 Simple Rules for Improving Your Standardized Fields and Terms

Oct 21, 2025
RC
Rhiannon Cameron
🏛️ Simon Fraser University

Scientific data often suffers from poor discoverability, limited sharing, inefficient reuse, and high curation costs due to inadequate standardization of fields and terminology. To address these challenges, this paper proposes a FAIR-aligned standardization framework. Methodologically, it integrates structured vocabulary design, context-aware metadata modeling, and data homogenization strategies to systematically mitigate semantic noise and concept explosion. Crucially, it embeds ten actionable, principle-based rules into a dynamic, evolving data governance process. Empirical evaluation demonstrates that the framework significantly improves metadata quality and semantic consistency, reduces data management overhead, and enhances data findability, interoperability, and long-term reusability—thereby enabling robust, real-world implementation of the FAIR principles in scientific research settings.

Addressing challenges in standardizing research metadata vocabulariesOffering practical rules for FAIR-compliant metadata designProviding strategies to improve data findability and reusability

Leveraging Taxonomy and LLMs for Improved Multimodal Hierarchical Classification

Jan 12, 2025
SC
Shijing Chen
🏛️ University of New South Wales | Deakin University | University of Southampton | Macquarie University | Technology Innovation Institute | MBZUAI

Existing multi-level hierarchical classification (MLHC) methods often neglect parent-child category relationships, leading to predictions that violate hierarchical constraints. To address this, we propose a taxonomy-structured, transition-based cross-modal product classification framework. Our approach introduces a novel taxonomy-embedded transition mechanism that explicitly models hierarchical dependencies to ensure inter-layer prediction consistency. It integrates category-tree encoding, hierarchical prompt engineering, and multimodal feature alignment, while leveraging frozen large language models (LLMs) with lightweight adapters for end-to-end training—enabling LLM-agnostic deployment. Evaluated on the MEP-3M dataset, our method achieves a 23.6% improvement in hierarchical consistency and a 9.4% gain in overall accuracy over conventional LLM-based baselines. The framework effectively balances structural awareness with generalization capability, advancing robust, constraint-aware hierarchical classification.

Classification Rule ViolationIgnored Hierarchical RelationshipsMulti-Level Hierarchical Classification

Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models

Apr 08, 2024
SS
Sowmya S. Sundaram
🏛️ Stanford University

This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.

Enhance metadata standards adherenceImprove metadata curation automationIntegrate structured knowledge with LLMs

Taxonomy Tree Generation from Citation Graph

Oct 02, 2024
YH
Yuntong Hu
🏛️ Emory University | Shanghai University

Manually constructing disciplinary taxonomy trees for scientific knowledge organization is labor-intensive, prone to human bias, and often fails to incorporate low-citation yet high-impact publications. Method: We propose HiGTL, an end-to-end framework that jointly optimizes structural consistency and semantic coherence by integrating hierarchical graph clustering with LLM-driven, layer-wise concept concretization. It combines hierarchical graph clustering, LLM-based node concept generation, multi-objective joint fine-tuning, and text-citation dual-modality representation learning. Results: Experiments demonstrate that HiGTL-generated taxonomies significantly outperform state-of-the-art methods in hierarchical plausibility, concept accuracy, and cross-layer consistency. The framework enables interpretable, interactive, and human-guided taxonomy construction, effectively supporting systematic literature reviews and emerging trend identification.

Automates taxonomy generation from citation graphsGenerates coherent taxonomies with hierarchical concept verbalizationRecursively clusters papers using text and citations

Latest Papers

What's happening recently
View more

This study addresses the challenge of aligning and interpreting knowledge structures across heterogeneous textual corpora by proposing a term-centric, hierarchical knowledge construction framework. Departing from conventional full-document representations, the approach maps multi-source documents into a shared semantic space through automated term extraction and integrates domain priors with data-driven clustering to generate interpretable knowledge hierarchies. Evaluated on a newly introduced benchmark comprising over one million English–German multi-source documents, the method significantly improves cross-source consistency and hierarchy quality. Its practical utility is further demonstrated through the successful construction of regional innovation technology maps for Germany, highlighting its applicability in policy analysis and innovation monitoring.

cross-source alignmentheterogeneous corporaknowledge organization

Existing scientific classification systems struggle to capture fine-grained conceptual structures in research, while author-provided keywords—though specific—are often fragmented, redundant, and inconsistent in terminology. This work proposes a novel approach that integrates scientific text embeddings, large language models (LLMs), and graph-based community detection to construct an interpretable Concept layer beneath the Topic hierarchy of OpenAlex. By aggregating semantically related keywords into unified, reusable conceptual units and situating them within disciplinary hierarchies, this method enables a systematic transformation from heterogeneous terms to structured knowledge entities. It represents the first large-scale academic taxonomy to incorporate an LLM- and embedding-driven intermediate conceptual layer, substantially enhancing the precision and scalability of scholarly document organization, science mapping, research trend monitoring, and ontology construction.

author keywordsconceptual structurefine-grained taxonomy

This work addresses fine-grained hierarchical classification in regulatory-intensive domains such as customs tariff codes and export controls, where predictions must strictly adhere to hierarchical structures and rule-based boundaries. Existing approaches struggle to simultaneously ensure hierarchical validity, rule consistency, and boundary-aware reasoning. To bridge this gap, the paper formalizes, for the first time, a regulation-driven hierarchical classification task and introduces a constraint-aware hierarchical search framework. The method parses regulatory documents into a searchable tree and, at each step, retrieves only legally permissible candidate nodes, guiding path decisions through structured fields and evidential text snippets. Evaluated on four expert-validated datasets, the approach significantly outperforms baselines in average accuracy—particularly excelling in distinguishing adjacent categories and handling boundary cases—while producing auditable and traceable decision paths.

constraint-aware reasoningfine-grained classificationhierarchical text classification

This work proposes an interactive ontology construction paradigm that bridges the gap between purely manual and fully automated approaches, which are often hindered by laborious processes or insufficient user control, respectively. By leveraging weighted self-organizing maps, the method enables progressive clustering of tabular data while integrating instance grouping with mechanisms for defining conceptual intensions. This approach empowers users to flexibly adjust both the number of clusters and their semantic interpretations, thereby preserving the efficiency of automation while significantly enhancing controllability. As a result, it facilitates interpretable clustering of entities and the generation of high-quality ontological classifications directly from tabular data.

cluster analysisconcept identificationinteractive construction

Hot Scholars

PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
JL

Juntao Li

Soochow University
Language ModelsText Generation
XH

Xuming Hu

Assistant Professor, HKUST(GZ) / HKUST
Natural Language ProcessingLarge Language Model
FD

Franck Dernoncourt

NLP/ML Researcher. MIT PhD.
Machine LearningNeural NetworksNatural Language Processing