Score
Designs and implements classification and object-detection models and training procedures that combine limited labeled examples with abundant unlabeled data, employing techniques such as clustering, pseudo-labeling, consistency regularization, or teacher–student approaches. This includes building semi-supervised training pipelines (e.g., SSOD), loss formulations and evaluation workflows to measure classification and detection performance while reducing reliance on large labeled datasets.
This study addresses the challenge of effectively leveraging large-scale unlabeled images to improve object detection performance under limited annotation budgets. We systematically evaluate three prominent semi-supervised object detection methods—MixPL, Semi-DETR, and Consistent-Teacher—across MS-COCO, Pascal VOC, and a custom Beetle dataset, analyzing their trade-offs among accuracy, model size, and inference latency under varying labeling ratios. For the first time, we reveal consistent patterns of performance degradation as labeled data decreases, both on general-purpose and domain-specific datasets. Our empirical findings provide actionable insights and practical guidance for selecting appropriate semi-supervised approaches in resource-constrained scenarios.
Deep visual models heavily rely on large-scale annotated data, hindering their deployment in low-label-resource scenarios. Method: This paper presents a unified survey of pseudo-labeling techniques across semi-supervised, self-supervised, and unsupervised learning. We propose, for the first time, a cross-paradigm pseudo-labeling conceptual framework that identifies methodological commonalities—namely label generation, confidence-based filtering, consistency regularization, and dynamic weighting—across these paradigms. We further introduce curriculum learning strategies and self-supervised regularization mechanisms to enable synergistic optimization among paradigms. Contribution/Results: We establish a comprehensive taxonomy covering pseudo-label generation, filtering, regularization, and cross-paradigm transfer; clarify the technical evolution trajectory; and empirically validate the feasibility of cross-paradigm pseudo-label transfer. Our work provides both theoretical foundations and reproducible practical paradigms for developing vision models with minimal annotation cost.
This work addresses the limitation of pseudo-label quality in semi-supervised image classification by proposing a contrastive learning framework integrated with a distribution matching mechanism. For the first time in semi-supervised contrastive learning, the method explicitly aligns the feature distributions of labeled and unlabeled data by minimizing the divergence in their statistical characteristics within the embedding space, thereby enhancing the reliability of pseudo-labels and the model’s generalization capability. Extensive experiments on multiple standard image classification benchmarks demonstrate that the proposed approach significantly outperforms existing state-of-the-art methods, confirming the effectiveness and novelty of incorporating feature distribution alignment to improve semi-supervised learning performance.
Oriented object detection in aerial imagery suffers from high annotation costs, and existing semi-supervised methods are limited to axis-aligned bounding boxes. Method: This paper pioneers the extension of semi-supervised learning to oriented object detection, proposing three novel components: (1) instance-aware dense sampling (SIDS), (2) geometry-aware adaptive weighting (GAW) loss, and (3) noise-driven global consistency (NGC) regularization. These jointly model the many-to-many set-level relationship between pseudo-labels and predictions, integrating geometric feature modeling with global layout constraints. Contribution/Results: On DOTA-V1.5 and DOTA-V2.0, our method achieves state-of-the-art (SOTA) performance using only 10%–30% labeled data—outperforming prior SOTA by 2.14–2.90 mAP. Under full supervision, it attains 72.48 mAP, surpassing previous SOTA by +1.82 mAP. Moreover, the framework generalizes effectively to diverse oriented detectors and multi-view 3D detection architectures.
Existing semi-supervised meta-training (SSMT) methods rely on class-aware sampling of unlabeled data, contradicting their fundamental “no class labels” premise and hindering few-shot learning (FSL) in realistic low-resource settings. To address this, we propose the first truly class-agnostic SSMT paradigm: Pseudo-Label-driven Meta-Learning (PLML). PLML bridges labeled and unlabeled data via adaptive pseudo-labeling, integrates feature-smoothing regularization, and employs noise-robust fine-tuning—eliminating explicit assumptions about class structure. Compatible with mainstream meta-learners (e.g., MAML, ProtoNet), PLML achieves significant improvements over prior SSMT methods on miniImageNet and CUB, mitigating performance degradation under extremely limited labeled data. Moreover, PLML reciprocally enhances the few-shot generalization of diverse self-supervised learning (SSL) algorithms. Our work establishes a principled foundation for label-efficient, class-agnostic meta-learning.
This work addresses the performance bottlenecks in semi-supervised semantic segmentation caused by noisy pseudo-labels and domain discrepancies between labeled and unlabeled data. To mitigate these issues, the authors propose a novel approach that integrates ClassMix augmentation with supervised–unsupervised feature alignment. Specifically, ground-truth class regions from labeled images are pasted onto unlabeled images and their corresponding pseudo-labels, while a feature discriminator enforces alignment of the model’s predictions with the feature distribution of the labeled data. This dual strategy effectively reduces the adverse effects of inaccurate pseudo-labels and domain shift. Experimental results on the CHASE and COVID-19 datasets demonstrate a consistent improvement, with an average mIoU gain of 2.07%, substantially outperforming existing semi-supervised segmentation methods.
To address low pseudo-label quality and poor generalization in few-shot semi-supervised text classification, this paper proposes a novel framework integrating target-masked unsupervised pretraining with teacher-student collaborative learning. Methodologically, it introduces target masking—previously unexplored in pseudo-labeling pretraining—to explicitly model class-distribution priors, thereby enhancing pseudo-label reliability and cross-lingual transferability. It further combines dynamic-threshold pseudo-label generation with bilingual consistency regularization over English and Swedish. Evaluated on three low-resource text classification benchmarks, the approach significantly outperforms baselines including Meta Pseudo Labels: under extreme few-shot settings (16–64 gold-labeled examples per class), it achieves up to a 4.2% absolute accuracy improvement. Results demonstrate superior effectiveness and robustness in ultra-low-resource scenarios, validating both the design rationale and practical utility of the proposed framework.
To address the high annotation cost of oriented object detection (OOD) under weak supervision, this paper proposes PWOOD, a semi-weakly supervised framework that leverages only partial weak annotations—such as axis-aligned bounding boxes or single points—together with abundant unlabeled data. Methodologically, PWOOD introduces an OS-Student model that explicitly encodes orientation and scale information, and incorporates a class-agnostic pseudo-label filtering (CPF) strategy to eliminate reliance on fixed confidence thresholds. It adopts a student–teacher paradigm integrating consistency regularization and dynamic pseudo-label refinement. On DOTA and DIOR benchmarks, PWOOD achieves performance comparable to or surpassing state-of-the-art semi-supervised methods, while requiring significantly less annotation effort than existing weakly supervised approaches. To our knowledge, PWOOD is the first work to systematically establish a partial weak supervision paradigm for OOD, bridging the gap between semi-supervised learning and practical annotation constraints in remote sensing and aerial imagery analysis.
This work proposes a self-supervised feature learning method specifically designed for object detection to address the heavy reliance on large-scale annotated data. By pretraining the feature extractor on unlabeled data and guiding the model to focus on semantically informative object regions, the approach significantly enhances the representational capacity of the detector under limited annotation budgets. Experimental results demonstrate that the proposed method outperforms conventional ImageNet-pretrained models across multiple object detection benchmarks, achieving not only improved detection accuracy but also greater robustness and reliability.
Existing semi-supervised 3D object detection methods suffer from low-quality pseudo-labels, manually tuned confidence thresholds, and insufficient exploitation of contextual information. Method: We propose an adaptive pseudo-label selection framework within a teacher–student paradigm. It introduces a learnable, context-aware thresholding module that dynamically generates class- and distance-dependent confidence thresholds based on object proximity, category, and model learning status. A soft supervision strategy is incorporated to mitigate noise from erroneous pseudo-labels. Furthermore, dual-network collaboration enables score fusion and quality assessment, with spatial alignment between pseudo-labels and ground-truth bounding boxes serving as the supervision signal. Results: Evaluated on KITTI and Waymo Open Dataset, our method achieves significant improvements in detection accuracy and recall, particularly for hard examples, and consistently outperforms state-of-the-art semi-supervised 3D detectors.