Score
Design and build structured classification systems and controlled vocabularies—hierarchical or flat—that define coherent categories and relations for tasks, datasets, labels, attributes, interactions, systems, evaluations, conceptual constructs, and sources-of-authority. Specify category definitions, boundaries, metadata and mapping rules, produce annotation guidelines and item-to-category mappings, and use the taxonomy to enable organized analysis, benchmarking, gap identification, and unified cross-study comparisons.
This work addresses the challenge of efficiently constructing a comprehensive and well-structured taxonomy of artificial intelligence skills and tasks from massive hiring data. To this end, the authors propose TaxonomyBuilder, a framework that integrates systematic data filtering, clustering algorithms, and large language model–enhanced hierarchical label generation to automatically derive domain-specific taxonomies from curated, high-quality data subsets. Experimental results demonstrate that taxonomies built from filtered data exhibit significantly broader coverage and superior structural coherence compared to those generated from raw, unfiltered data using existing methods. The study thus establishes a novel paradigm for data-driven, automated taxonomy construction in specialized domains.
Scientific data often suffers from poor discoverability, limited sharing, inefficient reuse, and high curation costs due to inadequate standardization of fields and terminology. To address these challenges, this paper proposes a FAIR-aligned standardization framework. Methodologically, it integrates structured vocabulary design, context-aware metadata modeling, and data homogenization strategies to systematically mitigate semantic noise and concept explosion. Crucially, it embeds ten actionable, principle-based rules into a dynamic, evolving data governance process. Empirical evaluation demonstrates that the framework significantly improves metadata quality and semantic consistency, reduces data management overhead, and enhances data findability, interoperability, and long-term reusability—thereby enabling robust, real-world implementation of the FAIR principles in scientific research settings.
Existing multi-level hierarchical classification (MLHC) methods often neglect parent-child category relationships, leading to predictions that violate hierarchical constraints. To address this, we propose a taxonomy-structured, transition-based cross-modal product classification framework. Our approach introduces a novel taxonomy-embedded transition mechanism that explicitly models hierarchical dependencies to ensure inter-layer prediction consistency. It integrates category-tree encoding, hierarchical prompt engineering, and multimodal feature alignment, while leveraging frozen large language models (LLMs) with lightweight adapters for end-to-end training—enabling LLM-agnostic deployment. Evaluated on the MEP-3M dataset, our method achieves a 23.6% improvement in hierarchical consistency and a 9.4% gain in overall accuracy over conventional LLM-based baselines. The framework effectively balances structural awareness with generalization capability, advancing robust, constraint-aware hierarchical classification.
This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.
Manually constructing disciplinary taxonomy trees for scientific knowledge organization is labor-intensive, prone to human bias, and often fails to incorporate low-citation yet high-impact publications. Method: We propose HiGTL, an end-to-end framework that jointly optimizes structural consistency and semantic coherence by integrating hierarchical graph clustering with LLM-driven, layer-wise concept concretization. It combines hierarchical graph clustering, LLM-based node concept generation, multi-objective joint fine-tuning, and text-citation dual-modality representation learning. Results: Experiments demonstrate that HiGTL-generated taxonomies significantly outperform state-of-the-art methods in hierarchical plausibility, concept accuracy, and cross-layer consistency. The framework enables interpretable, interactive, and human-guided taxonomy construction, effectively supporting systematic literature reviews and emerging trend identification.
This study addresses the challenge of aligning and interpreting knowledge structures across heterogeneous textual corpora by proposing a term-centric, hierarchical knowledge construction framework. Departing from conventional full-document representations, the approach maps multi-source documents into a shared semantic space through automated term extraction and integrates domain priors with data-driven clustering to generate interpretable knowledge hierarchies. Evaluated on a newly introduced benchmark comprising over one million English–German multi-source documents, the method significantly improves cross-source consistency and hierarchy quality. Its practical utility is further demonstrated through the successful construction of regional innovation technology maps for Germany, highlighting its applicability in policy analysis and innovation monitoring.
Existing scientific classification systems struggle to capture fine-grained conceptual structures in research, while author-provided keywords—though specific—are often fragmented, redundant, and inconsistent in terminology. This work proposes a novel approach that integrates scientific text embeddings, large language models (LLMs), and graph-based community detection to construct an interpretable Concept layer beneath the Topic hierarchy of OpenAlex. By aggregating semantically related keywords into unified, reusable conceptual units and situating them within disciplinary hierarchies, this method enables a systematic transformation from heterogeneous terms to structured knowledge entities. It represents the first large-scale academic taxonomy to incorporate an LLM- and embedding-driven intermediate conceptual layer, substantially enhancing the precision and scalability of scholarly document organization, science mapping, research trend monitoring, and ontology construction.
This work addresses fine-grained hierarchical classification in regulatory-intensive domains such as customs tariff codes and export controls, where predictions must strictly adhere to hierarchical structures and rule-based boundaries. Existing approaches struggle to simultaneously ensure hierarchical validity, rule consistency, and boundary-aware reasoning. To bridge this gap, the paper formalizes, for the first time, a regulation-driven hierarchical classification task and introduces a constraint-aware hierarchical search framework. The method parses regulatory documents into a searchable tree and, at each step, retrieves only legally permissible candidate nodes, guiding path decisions through structured fields and evidential text snippets. Evaluated on four expert-validated datasets, the approach significantly outperforms baselines in average accuracy—particularly excelling in distinguishing adjacent categories and handling boundary cases—while producing auditable and traceable decision paths.
This work proposes an interactive ontology construction paradigm that bridges the gap between purely manual and fully automated approaches, which are often hindered by laborious processes or insufficient user control, respectively. By leveraging weighted self-organizing maps, the method enables progressive clustering of tabular data while integrating instance grouping with mechanisms for defining conceptual intensions. This approach empowers users to flexibly adjust both the number of clusters and their semantic interpretations, thereby preserving the efficiency of automation while significantly enhancing controllability. As a result, it facilitates interpretable clustering of entities and the generation of high-quality ontological classifications directly from tabular data.