Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning

📅 2026-03-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses two critical challenges in generative zero-shot learning: the class-instance gap caused by intra-class variation and the domain gap arising from misalignment between semantic and visual feature distributions. To tackle these issues, the authors propose a unified modeling and alignment framework. It incorporates an Attribute Distribution Modeling (ADM) module to learn transferable class-level attribute distributions and sample instance-level attributes, along with a Visual-Guided Alignment (VGA) module that leverages visual information to refine semantic representations and explicitly align semantic and visual spaces. The proposed method achieves significant performance gains, improving state-of-the-art results by 4.7% on AWA2 and 6.1% on SUN benchmarks. Furthermore, it functions as a plug-and-play component that can effectively enhance other generative ZSL models.

Technology Category

Computer Vision: Generative Adversarial Networks (GANs) for VisionMachine Learning: Transfer, Domain Adaptation, Multi-Task LearningNatural Language Processing: Generation

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
Generative zero-shot learning (ZSL) synthesizes features for unseen classes, leveraging semantic conditions to transfer knowledge from seen classes. However, it also introduces two intrinsic challenges: (1) class-level attributes fails to capture instance-specific visual appearances due to substantial intra-class variability, thus causing the class-instance gap; (2) the substantial mismatch between semantic and visual feature distributions, manifested in inter-class correlations, gives rise to the semantic-visual domain gap. To address these challenges, we propose an Attribute Distribution Modeling and Semantic-Visual Alignment (ADiVA) approach, jointly modeling attribute distributions and performing explicit semantic-visual alignment. Specifically, our ADiVA consists of two modules: an Attribute Distribution Modeling (ADM) module that learns a transferable attribute distribution for each class and samples instance-level attributes for unseen classes, and a Visual-Guided Alignment (VGA) module that refines semantic representations to better reflect visual structures. Experiments on three widely used benchmark datasets demonstrate that ADiVA significantly outperforms state-of-the-art methods (e.g., achieving gains of 4.7% and 6.1% on AWA2 and SUN, respectively). Moreover, our approach can serve as a plugin to enhance existing generative ZSL methods.
Problem

Research questions and friction points this paper is trying to address.

zero-shot learning
attribute distribution
semantic-visual alignment
intra-class variability
domain gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attribute Distribution Modeling
Semantic-Visual Alignment
Generative Zero-shot Learning
Instance-level Attributes
Visual-Guided Alignment
🔎 Similar Papers
H
Haojie Pu
School of Computer Science and Engineering, Southeast University, Nanjing, China
Z
Zhuoming Li
School of Computer Science and Engineering, Southeast University, Nanjing, China
Y
Yongbiao Gao
Qilu University of Technology (Shandong Academy of Sciences), Jinan, China
Y
Yuheng Jia
School of Computer Science and Engineering, Southeast University, Nanjing, China