Score
Designs and evaluates classification models or systems that can assign a label to an input after observing only a single labeled example of each target class; this entails building or selecting similarity metrics, embedding spaces, prototypes, or generative/augmentation mechanisms that enable accurate labeling from one-shot examples. Focuses on methods and evaluations that generalize to new, previously unseen classes without retraining the core model.
This study addresses the limitations of evaluating multiclass classifiers using single performance metrics, which often leads to misleading conclusions. To overcome this, the work proposes a multidimensional evaluation paradigm that leverages the PyCM library to construct a comprehensive analytical framework, enabling systematic comparison of classifier performance across a diverse set of evaluation metrics. Through two case studies, the research uncovers nuanced performance trade-offs that conventional metrics fail to capture, thereby demonstrating the necessity and effectiveness of multidimensional assessment in model selection and optimization. The findings further highlight the unique value of PyCM in facilitating thorough and precise evaluation of multiclass classification systems.
In scenarios with scarce labeled data, conventional classifier evaluation suffers from high bias and variance. To address this, we propose Semi-Supervised Model Evaluation (SSME), a framework that jointly models a small set of labeled instances and a large pool of unlabeled data by aggregating continuous prediction scores from multiple classifiers. SSME estimates the joint distribution of true labels and predictions, enabling unbiased estimation of standard evaluation metrics—including accuracy, F1, and calibration error—without requiring additional ground-truth annotations. This work introduces two key advances: (i) fine-grained subgroup evaluation under label scarcity, and (ii) principled evaluation of large language model (LLM) outputs, breaking the reliance on fully labeled test sets. Extensive experiments across four domains—healthcare diagnosis, content moderation, molecular property prediction, and image annotation—demonstrate that SSME reduces estimation error by 5.1× over supervised baselines and outperforms the best prior method by 2.4×, while substantially improving robustness for subgroup and LLM evaluations.
Deep visual models heavily rely on large-scale annotated data, hindering their deployment in low-label-resource scenarios. Method: This paper presents a unified survey of pseudo-labeling techniques across semi-supervised, self-supervised, and unsupervised learning. We propose, for the first time, a cross-paradigm pseudo-labeling conceptual framework that identifies methodological commonalities—namely label generation, confidence-based filtering, consistency regularization, and dynamic weighting—across these paradigms. We further introduce curriculum learning strategies and self-supervised regularization mechanisms to enable synergistic optimization among paradigms. Contribution/Results: We establish a comprehensive taxonomy covering pseudo-label generation, filtering, regularization, and cross-paradigm transfer; clarify the technical evolution trajectory; and empirically validate the feasibility of cross-paradigm pseudo-label transfer. Our work provides both theoretical foundations and reproducible practical paradigms for developing vision models with minimal annotation cost.
Evaluating and comparing multi-label classifiers in high-dimensional label spaces remains challenging due to the lack of intuitive, scalable visualization methods. To address this, we propose an interactive visual analytics framework that operates without reliance on confusion matrices. The framework jointly models predictions from three complementary perspectives—instances, labels, and classifiers—and integrates label-level performance aggregation, coordinated multi-view navigation, and scalable rendering for comparative analysis across multiple classifiers. Its key innovation lies in decoupling visualization design from the number of labels, thereby enabling real-time exploration even with hundreds of labels. A user study demonstrates that our approach significantly improves both assessment efficiency and analytical insight depth. It achieves high scalability and strong interpretability while preserving operational simplicity—making it particularly suitable for diagnosing classifier behavior in large-scale multi-label settings.
This work addresses the challenge faced by non-technical users in formulating accurate natural language category descriptions for open-vocabulary object detection. We propose the first iterative human-in-the-loop feedback mechanism specifically designed for refining such textual class descriptions. Our method integrates text embedding analysis with contrastive example embedding synthesis, enabling users to dynamically define novel categories and iteratively improve description quality *in situ*, without retraining the detector. Evaluated across multiple state-of-the-art open-vocabulary detectors—including GLIP and GroundingDINO—the approach consistently improves detection accuracy (mAP gains of +3.2–5.7), while ensuring interpretability of outputs. Our key contributions are: (1) the first integration of human-AI iterative feedback into textual prompt engineering for open-vocabulary detection; (2) a novel description optimization paradigm grounded in contrastive embedding synthesis; and (3) empirical validation of the mechanism’s cross-model generalizability and robustness.
In few-shot learning, existing metric-based meta-learning approaches suffer from degraded generalization to unseen classes due to over-reliance on deep metrics optimized for seen classes. To address this, we propose a meta-component composition framework that models classifiers as reconfigurable sets of meta-components. During meta-training, orthogonal regularization explicitly decouples these components, enhancing their diversity and functional specificity—thereby enabling effective extraction of task-invariant discriminative substructures. This decoupling mitigates overfitting to seen classes and improves cross-class generalization. Evaluated on standard benchmarks including Mini-ImageNet and Tiered-ImageNet, our method achieves significant improvements over state-of-the-art metric-learning approaches. Empirical results validate the efficacy of both meta-component decoupling and compositional modeling for robust few-shot classification.
This study addresses a critical limitation in existing unsupervised feature selection methods, which are predominantly evaluated under single-label settings—a practice prone to performance comparison bias due to the arbitrary choice of labels. To overcome this issue, the work systematically exposes the shortcomings of the conventional evaluation paradigm and proposes replacing it with a multi-label classification framework, thereby establishing a more equitable and reliable assessment protocol. Extensive cross-method and cross-dataset experiments conducted on 21 real-world multi-label datasets demonstrate that the relative performance rankings of feature selection algorithms shift substantially under the proposed paradigm, strongly validating the necessity and effectiveness of the multi-label evaluation strategy.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This work addresses the heavy reliance of CLIP-style models on extensive human annotations for fine-grained or complex categories. To reduce dependence on costly manual supervision, we propose a weak-model-supervised strong-model classification paradigm. Our core method, Class Prototype Learning (CPL), leverages pseudo-labels generated by a lightweight weak model to construct robust and transferable class prototypes, which are then refined via contrastive learning to enhance vision-language alignment. Crucially, CPL is the first approach to successfully extend weak-to-strong generalization to multimodal vision-language settings. Extensive experiments under constrained regimes—including few-shot classification and low-resource pretraining—demonstrate that CPL consistently outperforms strong baselines, achieving an average accuracy gain of 3.67% across benchmarks. The improvement is particularly pronounced under annotation scarcity, validating CPL’s efficacy in low-supervision scenarios.
This work addresses the challenge of efficiently and reliably evaluating newly released machine learning models on unlabeled data without incurring costly annotations or repeated fine-tuning. The authors propose MetaEvaluator, the first model-agnostic, label-free, and training-free evaluation framework that eliminates the need for per-model retraining. Leveraging meta-learning, MetaEvaluator learns a transferable evaluator initialization from a pool of reference models, enabling rapid unsupervised performance estimation for models of unseen architectures and modalities. Extensive experiments demonstrate that MetaEvaluator consistently and accurately predicts model performance across diverse datasets and model types, substantially reducing evaluation costs and facilitating large-scale benchmarking on unlabeled data.