hierarchy-aware contrastive learning

Designs and implements contrastive learning objectives and training pipelines that explicitly encode hierarchical relationships among labels or instances; this includes constructing multi-level positive and negative pairs, applying hierarchy-informed negative sampling, formulating hierarchical contrastive loss terms that enforce parent–child semantic consistency and coarse-to-fine alignment, and optionally integrating cluster-based procedures (for example HDBSCAN) to infer or refine hierarchical supervision and tighten boundaries between nearby subtypes.

hierarchy-awarecontrastivelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Multi-level Supervised Contrastive Learning

Feb 04, 2025
NG
Naghmeh Ghanooni
🏛️ RPTU Kaiserslautern-Landau | RIKEN

Existing contrastive learning approaches typically employ a single projection head, limiting their ability to model fine-grained semantic similarities among samples with multi-label and hierarchical label structures—especially under few-shot settings. To address this, we propose Multi-level Supervised Contrastive Learning (MSCL), a novel framework featuring multiple dedicated nonlinear projection heads, each explicitly designed to capture inter-label and inter-hierarchical similarities. MSCL introduces a hierarchy-aware positive/negative sample construction strategy and jointly optimizes multiple supervised contrastive losses. Notably, this is the first work to explicitly integrate multi-level semantic supervision into the contrastive learning architecture. Extensive experiments on text and image multi-label and hierarchical classification benchmarks demonstrate substantial improvements over state-of-the-art methods, with particularly pronounced gains under low-resource conditions—achieving average accuracy improvements of +3.2% to +5.8%.

Address limitations of single projection headEnhance similarity capture in contrastive learningImprove performance in multi-label classification

Feature Identification for Hierarchical Contrastive Learning

Oct 01, 2025
JO
Julius Ott
🏛️ Technical University Munich | Infineon Technologies AG | Friedrich-Alexander-Universität Erlangen-Nürnberg

Hierarchical classification often neglects inter-class structural relationships and suffers from severe class imbalance (long-tail distribution). Method: This paper proposes two novel hierarchical contrastive learning approaches that jointly model hierarchical structure and long-tail distributions within a unified contrastive framework—first of its kind. Specifically, it employs Gaussian Mixture Models (GMMs) to capture hierarchy-specific feature distributions and integrates attention mechanisms to explicitly encode cross-level semantic associations. The method enables fine-grained feature disentanglement and cross-level clustering, emulating human hierarchical cognition. Contribution/Results: Evaluated via linear probing on CIFAR-100 and ModelNet40, the proposed methods achieve accuracy improvements of 2 percentage points over current state-of-the-art methods, demonstrating superior effectiveness and generalization capability in hierarchical representation learning.

Addressing imbalanced class distribution in hierarchical learning tasksCapturing inter-class relationships across hierarchical classification levelsImproving feature identification for fine-grained hierarchical clustering

This work addresses the hierarchical inconsistency problem in multi-level visual classification, where fine-grained predictions often conflict with their parent categories. To mitigate this issue, the authors propose a hierarchy-constrained contrastive learning mechanism that performs contrastive optimization exclusively within the same semantic level, thereby eliminating interference from cross-level false negatives. Additionally, a group-balanced optimization strategy is introduced to ensure adequate training across all hierarchy levels. Built upon the BioCLIP framework, the method jointly optimizes representations in both Euclidean and hyperbolic spaces, significantly improving hierarchical consistency and classification performance. Evaluated on benchmarks including iNaturalist 2021, the approach achieves a 30.47% average improvement in cross-level accuracy over baseline methods and demonstrates notably enhanced consistency under zero-shot settings.

contrastive learningfine-grained vision classificationhierarchical classification

Medical image labels inherently exhibit hierarchical structure (e.g., organ → tissue → subtype), yet prevailing self-supervised learning (SSL) methods ignore this hierarchy, yielding semantically inconsistent representations. To address this, we propose the first hierarchy-aware contrastive learning framework compatible with both Euclidean and hyperbolic embeddings—without modifying network architecture. Our method leverages the label taxonomy as explicit supervisory signal and evaluation benchmark. Key contributions are: (1) a Hierarchical Weighted Contrastive (HWC) loss that modulates attraction/repulsion strength between positive/negative pairs based on path weights in the label tree; and (2) a Level-Aware Margin (LAM) mechanism that dynamically adjusts contrastive margins according to semantic granularity. We evaluate using hierarchy-sensitive metrics (e.g., HF1, H-Acc) across multiple medical imaging benchmarks, achieving significant improvements over state-of-the-art methods. Ablation studies confirm the efficacy of each component, demonstrating superior taxonomy alignment and enhanced semantic interpretability.

Current approaches lack evaluation metrics for hierarchy faithfulnessExisting methods fail to preserve hierarchical relationships in medical imagingStandard self-supervised learning ignores medical label taxonomy structure

SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery

Oct 21, 2025
ZH
Zhenqi He
🏛️ The University of Hong Kong

This paper addresses Generalized Category Discovery (GCD), the task of jointly identifying both known and unknown categories in images under partial supervision. Existing approaches suffer from limited generalizability due to reliance on single-level semantics or manually designed hierarchies. To overcome this, we propose SEAL, a Semantic-aware Hierarchical learning framework. Its key contributions are: (1) hierarchical semantic-guided soft contrastive learning, which leverages natural taxonomic relationships to generate informative soft negative samples; and (2) a cross-granularity consistency module that enforces alignment between fine-grained and coarse-grained predictions to improve semantic coherence. SEAL achieves state-of-the-art performance on fine-grained benchmarks—including SSB, Oxford-Pets, and Herbarium19—and demonstrates strong generalization to coarse-grained datasets, validating its robustness across semantic granularities.

Categorizing unlabeled images from known and unknown classesImproving generalization and scalability in category discoveryOvercoming limitations of single-level semantics and manual hierarchies

Latest Papers

What's happening recently
View more

This work addresses the limitation of traditional semi-supervised hierarchical clustering, which relies solely on leaf-node constraints and thus struggles to effectively guide the formation of non-leaf hierarchical structures, often yielding trees inconsistent with ground-truth hierarchies. To overcome this, the authors propose a novel semi-supervised hyperbolic hierarchical clustering method that leverages set-level structural priors. Specifically, leaf-level constraints are transformed into semantically coherent sample sets, which serve as soft priors for subtree levels to guide end-to-end tree optimization. The approach innovatively treats sets as fundamental modeling units and integrates constrained consistent embedding, set partitioning, inter-set similarity estimation, and continuous tree optimization in hyperbolic space to enable effective supervision of non-leaf structures. Evaluated on eleven benchmark datasets, the method significantly outperforms existing baselines in both label consistency and tree quality metrics.

hierarchical organizationleaf-level supervisionsemi-supervised hierarchical clustering

This study investigates whether vision-language models (VLMs) can acquire concept-level abstraction capabilities without explicit high-level semantic supervision. To address this, we propose a group-wise contrastive learning framework tailored for the CLEAR GLASS model. Leveraging our self-constructed MAGIC dataset—comprising semantically grouped image-text pairs—we introduce a novel group-wise contrastive loss that jointly optimizes inter-group discrimination and intra-group alignment, thereby inducing concept-level semantic representations in the latent space. Our method integrates implicit semantic space modeling with unsupervised group induction, eliminating reliance on manually annotated high-level concept labels. Experiments demonstrate that the proposed approach significantly outperforms existing state-of-the-art methods on abstract concept recognition tasks. Notably, it is the first to enable VLMs to consistently and transferably emerge concept abstraction capabilities without any explicit high-level supervisory signals.

Developing strategies for encoding higher-level image conceptsEnhancing abstract concept recognition through novel contrastive learningEvaluating vision-language models' concept abstraction capacity

To address the challenges of preserving structural consistency in hierarchical multi-label classification (HMC) and mitigating loss-weight imbalance in multi-task learning (MTL), this paper proposes HCAL, a novel HMC classifier. First, it introduces a semantically consistent hierarchical feature aggregation mechanism to strengthen semantic correlations between parent and child labels. Second, it incorporates an adaptive task weighting strategy to dynamically alleviate optimization bias caused by “one-dominant, many-weak” label distributions. Third, it proposes prototype perturbation augmentation—injecting controlled noise into label prototypes—to enhance decision-boundary robustness, and defines Hierarchical Violation Rate (HVR) as a quantitative metric for structural consistency. Experiments on three benchmark datasets demonstrate that HCAL consistently outperforms state-of-the-art baselines, achieving average improvements of 2.1–4.7 percentage points in classification accuracy and reducing HVR by 18.3–32.6%, thereby validating its superior generalizability, structural consistency, and robustness.

Balancing loss weighting in multi-task learning frameworksMaintaining structural consistency in hierarchical multi-label classificationResolving one-strong-many-weak optimization bias in MTL

This work addresses the limitations of conventional multimodal representation learning, which relies on a shared-private dichotomy and struggles to capture latent factors shared only among subsets of modalities, often leading to excessive alignment of irrelevant signals and loss of complementary information. To overcome this, the authors propose a Hierarchical Contrastive Learning (HCL) framework that introduces, for the first time, a hierarchical latent variable structure to explicitly model globally shared, partially shared, and modality-specific components. A structure-aware contrastive objective is designed to align only those factors that are genuinely shared. Theoretical analysis establishes identifiability of the model without requiring correlation assumptions and provides recovery guarantees for the loading matrices along with bounds on prediction risk. Experiments demonstrate that HCL accurately recovers the hierarchical structure, effectively selects task-relevant components, and significantly improves representation quality and downstream prediction performance on multimodal electronic health records.

hierarchical structurelatent factorsmultimodal representation learning

This work addresses the challenge of distinguishing semantically similar sibling categories in few-shot hierarchical text classification, a problem exacerbated by insufficient domain knowledge. To this end, the authors propose a novel approach that integrates hierarchical knowledge-aware prompt tuning with sibling contrastive learning. Specifically, a hierarchical knowledge extraction module is designed to explicitly model hierarchical semantics, while a fine-grained contrastive learning mechanism is introduced among sibling categories to enhance the model’s discriminative capacity for easily confused classes. Notably, this method is the first to specifically target the differentiation of deep-level sibling categories. Experimental results on three benchmark datasets demonstrate substantial improvements over current state-of-the-art methods, effectively advancing performance in few-shot hierarchical text classification.

Data ScarcityFew-shot Hierarchical Text ClassificationLabel Hierarchy

Hot Scholars

MJ

Mostafa Jahanifar

AI Researcher
Multi-Modal Deep learningComputational PathologyGenerative AI
XY

Xulei Yang

Principal Scientist & Group Leader, A*STAR, Singapore
3D VisionArtificial IntelligenceMedical Imaging
SB

Seunghun Baek

GSAI, POSTECH, S. Korea
Machine LearningMedical ImagingComputer Vision
WL

Wang Lin

Zhejiang University
Computer VisionMulti-Modal LearningVideo Understanding
SY

Si Yong Yeo

Asst. Professor, Nanyang Technological University
Computer VisionMedical InformaticsArtificial IntelligenceMedical Imaging