harmonize hierarchical labels

Designs and implements artifacts and processes for aligning hierarchical label systems — e.g., multi‑level taxonomies, ontologies, mapping tables, and automated alignment algorithms — that reconcile disparate label schemes and levels of semantic granularity. Builds rules, transformation procedures, and validation checks to standardize names, merge or split classes across hierarchies, and ensure consistent labeling across systems, regions, and datasets.

harmonizehierarchicallabels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In domain ontology construction, mapping multi-source terminology to foundational concepts faces three key challenges: high cost and subjectivity of manual approaches, shallow semantic modeling and poor cross-domain consistency of automated methods, and weak interpretability. To address these, this paper proposes the first LLM-driven framework integrating expert calibration with iterative prompt optimization. The framework combines expert-guided annotation, multi-stage prompt engineering, and a human-in-the-loop validation cycle to generate concept links with high confidence and full interpretability. Evaluated on the concept necessity mapping task, it achieves an F1-score of 0.97—substantially surpassing the human baseline (0.68)—and marks the first instance of scalable ontology alignment that simultaneously attains expert-level accuracy and transparent, auditable reasoning.

Automating taxonomy alignment to replace costly manual expert reviewEnhancing transparency and consistency in automated concept mapping processesImproving semantic relationship handling in cross-domain taxonomy mapping

This work proposes Taxon, a framework designed to address the compliance risks arising from inaccurate matching between e-commerce products and multi-level national tax classification systems. Taxon integrates a multimodal feature-gating mixture-of-experts (MoE) architecture, a semantic consistency verification mechanism guided by distillation from large language models, and a multi-source training pipeline that unifies tax code databases, invoice logs, and merchant registration data. To enhance structural consistency, the framework further incorporates full hierarchical path reconstruction. Evaluated on both a newly curated TaxCode dataset and public benchmarks, Taxon achieves state-of-the-art performance and has been deployed in Alibaba’s tax system, handling over 500,000 queries daily while significantly improving accuracy, interpretability, and robustness.

e-commerce compliancehierarchical taxonomymulti-level classification

This study addresses the challenges posed by the high heterogeneity of healthcare data and the lack of effective metadata management, which often degrade conventional data lakes into “data swamps,” impeding data interoperability and machine learning (ML) readiness. To overcome these limitations, the authors propose a dual-hybrid semantic data lake architecture that synergistically integrates the dynamic modeling capabilities of knowledge graphs with the metadata generation power of large language models (LLMs). A human-in-the-loop validation mechanism is incorporated to enable automated metadata annotation and high-level semantic alignment. This approach establishes, for the first time, semantic linkages within a data lake explicitly oriented toward ML operability, substantially enhancing the discoverability and computability of heterogeneous medical data while supporting intelligent recommendation of suitable ML methods.

data lakeheterogeneous datainteroperability

Latest Papers

What's happening recently
View more

This work addresses fine-grained hierarchical classification in regulatory-intensive domains such as customs tariff codes and export controls, where predictions must strictly adhere to hierarchical structures and rule-based boundaries. Existing approaches struggle to simultaneously ensure hierarchical validity, rule consistency, and boundary-aware reasoning. To bridge this gap, the paper formalizes, for the first time, a regulation-driven hierarchical classification task and introduces a constraint-aware hierarchical search framework. The method parses regulatory documents into a searchable tree and, at each step, retrieves only legally permissible candidate nodes, guiding path decisions through structured fields and evidential text snippets. Evaluated on four expert-validated datasets, the approach significantly outperforms baselines in average accuracy—particularly excelling in distinguishing adjacent categories and handling boundary cases—while producing auditable and traceable decision paths.

constraint-aware reasoningfine-grained classificationhierarchical text classification

Existing approaches to scientific knowledge classification often suffer from semantic inconsistency and structural misalignment within hierarchical taxonomies, hindering their ability to effectively organize the rapidly expanding body of scholarly literature. This work proposes a hierarchical classification framework grounded in large language models, which integrates a bidirectional title generation mechanism—combining bottom-up abstraction with top-down constraints—to jointly optimize vertical alignment across levels and horizontal semantic coherence among sibling nodes. The method explicitly models semantic dependencies among nodes at the same hierarchy level, significantly enhancing the logical structure, semantic fidelity, and title quality of the resulting taxonomy. Extensive experiments demonstrate strong performance across multiple benchmark datasets and reveal robust cross-lingual generalization capabilities, particularly on Chinese scientific literature.

hierarchical structurelarge language modelsscientific literature

This study addresses the challenge that language models in high-stakes professional settings often encounter conflicting demands among user instructions, institutional authority, and domain-specific ethical norms, with implicit priority hierarchies potentially leading to harmful outputs that violate professional standards. Through systematic evaluation across 7,136 adversarial scenarios in legal and medical domains, the work assesses ten state-of-the-art models’ adherence to professional norms under both task-execution and advisory frameworks. It reveals, for the first time, an unstable hierarchy of alignment priorities and identifies “knowledge omission”—the suppression of critical facts in model outputs despite their correct internal recognition—as a core mechanism underlying norm violations. The findings demonstrate that prevailing models frequently breach professional standards, with alignment behavior exhibiting significant inconsistency across domains, tasks, and model architectures, underscoring the lack of robustness in current alignment approaches for high-risk applications.

alignmenthigh-stakes settingslanguage models

This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.

data curationentity identityentity resolution

Existing data mixing approaches are constrained by fixed semantic labels and a single granularity level, limiting their ability to flexibly explore combinatorial effects across multiple granularities. This work proposes a hierarchical, data-driven multi-granularity annotation framework that leverages learnable semantic transformations and a three-stage residual vector quantization scheme to generate up to 130,000 reusable hierarchical document codes, enabling dynamic navigation from coarse to fine granularities. Evaluated in a pretraining setting with 1B parameters and 25B tokens, the method—combined with an equal-subbucket coverage strategy—achieves an average performance gain of +0.0253 across 16 tasks at specific granularities, demonstrating a significant interaction effect between granularity selection and mixing strategy.

data mixinglabel granularitymixture design

Hot Scholars

YJ

Yueming Jin

Assistant Professor, National University of Singapore
Medical Image AnalysisSurgical AI&RoboticsMultimodal Learning
ZZ

Zhitao Zeng

National University of Singapore
Vision-Language Models
CH

Chang Han Low

National University of Singapore
Surgical AIMedical ImagingMulti-Agent SystemMulti-Modal
ZZ

Zhu Zhuo

National University of Singapore
Surgical Data ScienceMultimodal Large Language Model
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion