semi-supervised classification

Designs and implements classification and object-detection models and training procedures that combine limited labeled examples with abundant unlabeled data, employing techniques such as clustering, pseudo-labeling, consistency regularization, or teacher–student approaches. This includes building semi-supervised training pipelines (e.g., SSOD), loss formulations and evaluation workflows to measure classification and detection performance while reducing reliance on large labeled datasets.

semi-supervisedclassification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of effectively leveraging large-scale unlabeled images to improve object detection performance under limited annotation budgets. We systematically evaluate three prominent semi-supervised object detection methods—MixPL, Semi-DETR, and Consistent-Teacher—across MS-COCO, Pascal VOC, and a custom Beetle dataset, analyzing their trade-offs among accuracy, model size, and inference latency under varying labeling ratios. For the first time, we reveal consistent patterns of performance degradation as labeled data decreases, both on general-purpose and domain-specific datasets. Our empirical findings provide actionable insights and practical guidance for selecting appropriate semi-supervised approaches in resource-constrained scenarios.

data-scarce learningfew-shot learningobject detection

A Review of Pseudo-Labeling for Computer Vision

Aug 13, 2024
PK
Patrick Kage
🏛️ The University of Edinburgh | The University of Oklahoma

Deep visual models heavily rely on large-scale annotated data, hindering their deployment in low-label-resource scenarios. Method: This paper presents a unified survey of pseudo-labeling techniques across semi-supervised, self-supervised, and unsupervised learning. We propose, for the first time, a cross-paradigm pseudo-labeling conceptual framework that identifies methodological commonalities—namely label generation, confidence-based filtering, consistency regularization, and dynamic weighting—across these paradigms. We further introduce curriculum learning strategies and self-supervised regularization mechanisms to enable synergistic optimization among paradigms. Contribution/Results: We establish a comprehensive taxonomy covering pseudo-label generation, filtering, regularization, and cross-paradigm transfer; clarify the technical evolution trajectory; and empirically validate the feasibility of cross-paradigm pseudo-label transfer. Our work provides both theoretical foundations and reproducible practical paradigms for developing vision models with minimal annotation cost.

Connecting advancements in pseudo-labeling across different learning methodsExploring pseudo-labeling in semi-supervised and unsupervised learningReducing reliance on large labeled datasets for deep neural networks

This work addresses the limitation of pseudo-label quality in semi-supervised image classification by proposing a contrastive learning framework integrated with a distribution matching mechanism. For the first time in semi-supervised contrastive learning, the method explicitly aligns the feature distributions of labeled and unlabeled data by minimizing the divergence in their statistical characteristics within the embedding space, thereby enhancing the reliability of pseudo-labels and the model’s generalization capability. Extensive experiments on multiple standard image classification benchmarks demonstrate that the proposed approach significantly outperforms existing state-of-the-art methods, confirming the effectiveness and novelty of incorporating feature distribution alignment to improve semi-supervised learning performance.

contrastive learningdistribution matchingimage classification

SOOD++: Leveraging Unlabeled Data to Boost Oriented Object Detection

Jul 01, 2024
DL
Dingkang Liang
🏛️ Huazhong University of Science and Technology | Baidu Inc.

Oriented object detection in aerial imagery suffers from high annotation costs, and existing semi-supervised methods are limited to axis-aligned bounding boxes. Method: This paper pioneers the extension of semi-supervised learning to oriented object detection, proposing three novel components: (1) instance-aware dense sampling (SIDS), (2) geometry-aware adaptive weighting (GAW) loss, and (3) noise-driven global consistency (NGC) regularization. These jointly model the many-to-many set-level relationship between pseudo-labels and predictions, integrating geometric feature modeling with global layout constraints. Contribution/Results: On DOTA-V1.5 and DOTA-V2.0, our method achieves state-of-the-art (SOTA) performance using only 10%–30% labeled data—outperforming prior SOTA by 2.14–2.90 mAP. Under full supervision, it attains 72.48 mAP, surpassing previous SOTA by +1.82 mAP. Moreover, the framework generalizes effectively to diverse oriented detectors and multi-view 3D detection architectures.

Addressing arbitrary orientations and dense distribution in aerial objectsBoosting oriented object detection using unlabeled aerial image dataReducing high annotation costs for oriented objects in aerial imagery

Pseudo-Labeling Based Practical Semi-Supervised Meta-Training for Few-Shot Learning

Jul 14, 2022
XD
Xingping Dong
🏛️ Wuhan University | University of Chinese Academy of Sciences | United Arab Emirates University

Existing semi-supervised meta-training (SSMT) methods rely on class-aware sampling of unlabeled data, contradicting their fundamental “no class labels” premise and hindering few-shot learning (FSL) in realistic low-resource settings. To address this, we propose the first truly class-agnostic SSMT paradigm: Pseudo-Label-driven Meta-Learning (PLML). PLML bridges labeled and unlabeled data via adaptive pseudo-labeling, integrates feature-smoothing regularization, and employs noise-robust fine-tuning—eliminating explicit assumptions about class structure. Compatible with mainstream meta-learners (e.g., MAML, ProtoNet), PLML achieves significant improvements over prior SSMT methods on miniImageNet and CUB, mitigating performance degradation under extremely limited labeled data. Moreover, PLML reciprocally enhances the few-shot generalization of diverse self-supervised learning (SSL) algorithms. Our work establishes a principled foundation for label-efficient, class-agnostic meta-learning.

Enhancing FSL model performance under noise labels via pseudo-labeling frameworkOvercoming class-aware selection limits with truly unlabeled data utilizationReducing labeled data need in few-shot learning via semi-supervised meta-training

Latest Papers

What's happening recently
View more

This work addresses the performance bottlenecks in semi-supervised semantic segmentation caused by noisy pseudo-labels and domain discrepancies between labeled and unlabeled data. To mitigate these issues, the authors propose a novel approach that integrates ClassMix augmentation with supervised–unsupervised feature alignment. Specifically, ground-truth class regions from labeled images are pasted onto unlabeled images and their corresponding pseudo-labels, while a feature discriminator enforces alignment of the model’s predictions with the feature distribution of the labeled data. This dual strategy effectively reduces the adverse effects of inaccurate pseudo-labels and domain shift. Experimental results on the CHASE and COVID-19 datasets demonstrate a consistent improvement, with an average mIoU gain of 2.07%, substantially outperforming existing semi-supervised segmentation methods.

data quality gaplabel accuracypseudo-label noise

To address low pseudo-label quality and poor generalization in few-shot semi-supervised text classification, this paper proposes a novel framework integrating target-masked unsupervised pretraining with teacher-student collaborative learning. Methodologically, it introduces target masking—previously unexplored in pseudo-labeling pretraining—to explicitly model class-distribution priors, thereby enhancing pseudo-label reliability and cross-lingual transferability. It further combines dynamic-threshold pseudo-label generation with bilingual consistency regularization over English and Swedish. Evaluated on three low-resource text classification benchmarks, the approach significantly outperforms baselines including Meta Pseudo Labels: under extreme few-shot settings (16–64 gold-labeled examples per class), it achieves up to a 4.2% absolute accuracy improvement. Results demonstrate superior effectiveness and robustness in ultra-low-resource scenarios, validating both the design rationale and practical utility of the proposed framework.

Enhancing teacher-student model via objective masking pre-trainingEvaluating performance across multilingual datasets (English, Swedish)Improving semi-supervised text classification with few labeled examples

Partial Weakly-Supervised Oriented Object Detection

Jul 03, 2025
ML
Mingxin Liu
🏛️ Shanghai Jiao Tong University | Wuhan University | East China Normal University | Aerospace Information Research Institute | Southeast University

To address the high annotation cost of oriented object detection (OOD) under weak supervision, this paper proposes PWOOD, a semi-weakly supervised framework that leverages only partial weak annotations—such as axis-aligned bounding boxes or single points—together with abundant unlabeled data. Methodologically, PWOOD introduces an OS-Student model that explicitly encodes orientation and scale information, and incorporates a class-agnostic pseudo-label filtering (CPF) strategy to eliminate reliance on fixed confidence thresholds. It adopts a student–teacher paradigm integrating consistency regularization and dynamic pseudo-label refinement. On DOTA and DIOR benchmarks, PWOOD achieves performance comparable to or surpassing state-of-the-art semi-supervised methods, while requiring significantly less annotation effort than existing weakly supervised approaches. To our knowledge, PWOOD is the first work to systematically establish a partial weak supervision paradigm for OOD, bridging the gap between semi-supervised learning and practical annotation constraints in remote sensing and aerial imagery analysis.

Improving detection accuracy with partial weak supervisionLeveraging weak annotations for efficient model trainingReducing annotation cost in oriented object detection

This work proposes a self-supervised feature learning method specifically designed for object detection to address the heavy reliance on large-scale annotated data. By pretraining the feature extractor on unlabeled data and guiding the model to focus on semantically informative object regions, the approach significantly enhances the representational capacity of the detector under limited annotation budgets. Experimental results demonstrate that the proposed method outperforms conventional ImageNet-pretrained models across multiple object detection benchmarks, achieving not only improved detection accuracy but also greater robustness and reliability.

data annotationfeature representationlabeled data

Existing semi-supervised 3D object detection methods suffer from low-quality pseudo-labels, manually tuned confidence thresholds, and insufficient exploitation of contextual information. Method: We propose an adaptive pseudo-label selection framework within a teacher–student paradigm. It introduces a learnable, context-aware thresholding module that dynamically generates class- and distance-dependent confidence thresholds based on object proximity, category, and model learning status. A soft supervision strategy is incorporated to mitigate noise from erroneous pseudo-labels. Furthermore, dual-network collaboration enables score fusion and quality assessment, with spatial alignment between pseudo-labels and ground-truth bounding boxes serving as the supervision signal. Results: Evaluated on KITTI and Waymo Open Dataset, our method achieves significant improvements in detection accuracy and recall, particularly for hard examples, and consistently outperforms state-of-the-art semi-supervised 3D detectors.

Assessing pseudo-label quality using contextual information and learning statesOvercoming limitations of manual thresholding for pseudo-label selectionSelecting high-quality pseudo-labels in semi-supervised 3D object detection

Hot Scholars

TH

Thanh-Huy Nguyen

Carnegie Mellon University
Medical Image Analysis𝗖𝗼𝗺𝗽𝘂𝘁𝗲𝗿 𝗩𝗶𝘀𝗶𝗼𝗻Semi-Supervised Learning
BD

Bo Du

Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain
SY

Senthil Yogamani

Engineering Director, Data-centric AI for Autonomous Driving at Qualcomm Inc
Multimodal PerceptionData-centric AIIntelligent VehiclesAutonomous Driving
CC

Cornelia Caragea

University of Illinois at Chicago
Natural Language ProcessingDeep LearningInformation RetrievalArtificial Intelligence
JH

Junhui Hou

Department of Computer Science, City University of Hong Kong
Neural Spatial Computing