PrototypeFormer: Learning to Explore Prototype Relationships for Few-shot Image Classification

📅 2023-10-05
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient generalization in few-shot image classification caused by scarce support samples for novel classes, this paper proposes a Transformer-based prototypical relational modeling method. The approach integrates prototypical learning with meta-training without auxiliary modules. Its core innovations are: (i) the first explicit modeling of structured relationships among class prototypes using a Transformer architecture; and (ii) a lightweight, parameter-free contrastive learning mechanism that jointly optimizes prototype discriminability in an end-to-end manner. Evaluated on miniImageNet, the method achieves state-of-the-art accuracy of 97.07% (5-way 5-shot) and 90.88% (5-way 1-shot), surpassing prior art by 0.57% and 6.84%, respectively. These results significantly advance the performance frontier of few-shot classification.
📝 Abstract
Few-shot image classification has received considerable attention for overcoming the challenge of limited classification performance with limited samples in novel classes. Most existing works employ sophisticated learning strategies and feature learning modules to alleviate this challenge. In this paper, we propose a novel method called PrototypeFormer, exploring the relationships among category prototypes in the few-shot scenario. Specifically, we utilize a transformer architecture to build a prototype extraction module, aiming to extract class representations that are more discriminative for few-shot classification. Besides, during the model training process, we propose a contrastive learning-based optimization approach to optimize prototype features in few-shot learning scenarios. Despite its simplicity, our method performs remarkably well, with no bells and whistles. We have experimented with our approach on several popular few-shot image classification benchmark datasets, which shows that our method outperforms all current state-of-the-art methods. In particular, our method achieves 97.07% and 90.88% on 5-way 5-shot and 5-way 1-shot tasks of miniImageNet, which surpasses the state-of-the-art results with accuracy of 0.57% and 6.84%, respectively. The code will be released later.
Problem

Research questions and friction points this paper is trying to address.

Few-shot image classification challenge
Prototype relationships exploration
Transformer-based prototype extraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer for prototype extraction
Contrastive learning optimization
Few-shot classification enhancement
🔎 Similar Papers
Soochow University | Chinese Academy of Sciences | University of Chinese Academy of Sciences | Tsinghua University
F
Feihong He
School of Computer Science and Technology, Soochow University
G
Gang Li
Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Lingyu Si
Lingyu Si
Institute of Software Chinese Academy of Sciences
Computer visionmachine learingdeep learning
L
Leilei Yan
School of Computer Science and Technology, Soochow University
F
Fanzhang Li
School of Computer Science and Technology, Soochow University
F
Fuchun Sun
Department of Computer Science and Technology, Tsinghua University