Score
Designs, implements, and evaluates metric-based few-shot classification systems that represent each class by a prototype embedding computed from a small support set and classify query samples by their distance to these prototypes. This includes building prototypical networks and associated procedures for meta-training, rapid adaptation from a handful of examples, choice of embedding and distance functions, and prototype-based techniques to mitigate class imbalance.
Few-shot learning (FSL) addresses the generalization bottleneck of deep learning under data-scarce conditions. This work conducts a systematic literature review and paradigmatic reconstruction to unify FSL frameworks: it is the first to incorporate in-context learning into a coherent FSL taxonomy; introduces novel meta-learning subcategories—including neural processes and probabilistic meta-learning—to extend classical meta-learning theory; and establishes the first comprehensive FSL taxonomy encompassing supervised, semi-supervised, and unsupervised settings. Through cross-paradigm comparative analysis and application-domain mapping, we construct the most extensive FSL knowledge graph to date, clarifying core challenges—such as task distribution shift and prior modeling bias—and identifying key future directions, including enhanced interpretability and cross-modal transfer. The resulting framework provides both theoretical foundations and practical guidelines for algorithmic innovation and cross-domain deployment.
In few-shot learning, existing metric-based meta-learning approaches suffer from degraded generalization to unseen classes due to over-reliance on deep metrics optimized for seen classes. To address this, we propose a meta-component composition framework that models classifiers as reconfigurable sets of meta-components. During meta-training, orthogonal regularization explicitly decouples these components, enhancing their diversity and functional specificity—thereby enabling effective extraction of task-invariant discriminative substructures. This decoupling mitigates overfitting to seen classes and improves cross-class generalization. Evaluated on standard benchmarks including Mini-ImageNet and Tiered-ImageNet, our method achieves significant improvements over state-of-the-art metric-learning approaches. Empirical results validate the efficacy of both meta-component decoupling and compositional modeling for robust few-shot classification.
To address the limitations of single-scale feature matching in few-shot image classification—namely, loss of fine-grained details and insufficient generalization—this paper proposes a learnable multi-scale embedding framework. It employs a multi-output CNN to simultaneously extract shallow-level discriminative details and deep-level semantic features; introduces a staged self-attention mechanism to enhance cross-layer feature alignment; and integrates a learnable scale-weighting module for dynamic multi-scale fusion. Notably, this work is the first to jointly leverage staged self-attention and adaptive scale weighting for few-shot feature alignment, significantly improving prototype matching robustness and cross-domain generalization. The method achieves state-of-the-art performance on 5-way 1-shot and 5-shot tasks of MiniImageNet and FC100. Furthermore, extensive cross-domain evaluation across eight benchmark datasets demonstrates superior overall performance.
This paper addresses three key challenges in transductive few-shot learning (TFSL): inter-class confusion, embedding distribution bias, and the hubness problem. To tackle these, we propose an unbiased embedding classification framework. Methodologically: (1) we introduce a decentered covariance modeling strategy to mitigate centroid bias; (2) we design an adaptive nonlinear embedding optimization that jointly enforces local alignment and global uniformity; and (3) we develop a variational Sinkhorn classifier that jointly optimizes prototype distances and transductive clustering. Evaluated on standard TFSL benchmarks, our approach significantly outperforms state-of-the-art methods. Results validate the effectiveness of the “clustering-as-classification” paradigm—achieving robust embedding learning and high-accuracy classification with only a minimal number of labeled examples. This work offers a novel perspective for few-shot learning under low-resource settings.
To address insufficient generalization in few-shot image classification caused by scarce support samples for novel classes, this paper proposes a Transformer-based prototypical relational modeling method. The approach integrates prototypical learning with meta-training without auxiliary modules. Its core innovations are: (i) the first explicit modeling of structured relationships among class prototypes using a Transformer architecture; and (ii) a lightweight, parameter-free contrastive learning mechanism that jointly optimizes prototype discriminability in an end-to-end manner. Evaluated on miniImageNet, the method achieves state-of-the-art accuracy of 97.07% (5-way 5-shot) and 90.88% (5-way 1-shot), surpassing prior art by 0.57% and 6.84%, respectively. These results significantly advance the performance frontier of few-shot classification.
Generalized Category Discovery (GCD) aims to jointly cluster unlabeled data—containing both known and novel classes—by leveraging labeled data from known classes; its core challenge lies in accuracy imbalance caused by distributional ambiguity between known and novel classes. This paper proposes a Joint Prototype Learning (JPL) framework: (1) unifying prototype modeling for both known and novel classes to eliminate classifier bias; (2) introducing a two-level adaptive pseudo-labeling mechanism to mitigate confirmation bias; and (3) integrating contrastive regularization and clustering consistency constraints to align feature representations with clustering objectives, further enhanced by novel-class cardinality estimation and outlier detection for task-level co-optimization. Evaluated on both generic and fine-grained benchmarks, JPL achieves state-of-the-art performance, significantly improving balanced accuracy across known and novel classes while enhancing representation discriminability.
This work addresses the instability and performance limitations of few-shot learning during inference, which arise from the common assumption of batch-wise independence that prevents leveraging historical query samples. To overcome this, the authors propose the Incremental Prototype Enhancement Classifier (IPEC), which dynamically constructs an auxiliary set of high-confidence query samples and fuses it with the support set to progressively refine class prototypes. IPEC incorporates a dual-filtering mechanism that balances global confidence and local discriminability, along with a Bayesian-inspired prototype update strategy that treats the support set as a prior and the auxiliary set as likelihood-derived evidence. A two-stage “warm-up–test” inference protocol is introduced to move beyond static prototype representations. Extensive experiments demonstrate that IPEC significantly outperforms existing methods across multiple few-shot classification benchmarks, effectively enhancing both prototype stability and classification accuracy.
Zero-shot network content classification often suffers from semantic overlap and systematic misclassification due to ambiguous category definitions. This work identifies definition quality as a critical yet overlooked factor in zero-shot embedding systems and introduces the first training-free, iterative framework for refining category definitions. Leveraging large language models (LLMs) as feedback-driven optimizers, the approach dynamically refines semantic category prototypes—rather than model parameters—using structured signals derived from misclassified samples. We propose three LLM-guided refinement strategies: example-guided, confusion-aware, and history-aware refinement, and introduce B2MWT-10C, a new annotated benchmark comprising ten categories. Evaluated across 13 state-of-the-art embedding models, our method consistently improves classification performance, demonstrating that optimizing category definitions yields significant gains in zero-shot settings. The dataset and code are publicly released.
This paper addresses the challenge of fair comparison between meta-learning and full-class supervised training in unsupervised few-shot classification. To enable principled evaluation, we propose an entropy-constrained assessment framework that reveals meta-learning’s distinct advantages under low-entropy regimes, label noise, and task heterogeneity. We introduce MINO—a novel unsupervised meta-learning framework—featuring (i) a dynamic-head DBSCAN module for adaptive construction of unsupervised meta-tasks, and (ii) a stability-weighted meta-scaler to enhance robustness against label noise. Furthermore, we integrate entropy regularization with theoretical analysis to ensure interpretability and generalization guarantees. Extensive experiments on multiple unsupervised few-shot and zero-shot benchmarks demonstrate that MINO consistently outperforms state-of-the-art methods, especially under high label noise and task distribution shift. Our work establishes a new paradigm for applying meta-learning in realistic weakly supervised settings and provides strong empirical validation.
This work addresses the limitations of static prototypes in few-shot class-incremental learning, which are prone to representation bias from the backbone network and consequently hinder performance. To overcome this, the authors propose a novel paradigm that freezes the pre-trained feature extractor and instead fine-tunes learnable prototypes. They introduce a dual calibration mechanism—comprising class-specific and task-aware offset adjustments—that enables prototypes to dynamically adapt to new classes within a high-quality, fixed feature space. Remarkably, this approach requires only a minimal number of learnable parameters yet achieves substantial performance gains over existing methods across multiple benchmarks, significantly enhancing both discriminative capability and incremental learning efficacy.
Optimizing generalized evaluation metrics—such as Fβ, accuracy parity (AM), and Jaccard—under class imbalance and asymmetric misclassification costs remains challenging, as these metrics are non-decomposable and non-differentiable, precluding direct empirical risk minimization without probabilistic calibration or threshold tuning. Method: This paper introduces the first unified cost-sensitive learning framework that bypasses both probability estimation and threshold optimization. At its core lies a novel theory for constructing H-consistent surrogate losses with finite-sample generalization bounds. We further propose METRO, an algorithm that achieves provably optimal empirical risk minimization over arbitrary hypothesis classes. Contribution/Results: Evaluated on multiple multiclass imbalanced benchmarks, our method consistently outperforms state-of-the-art approaches. Empirical results validate its superior convergence behavior and generalization performance, demonstrating that theoretical guarantees translate effectively into practical gains.