Self-Supervised Federated Learning under Data Heterogeneity for Label-Scarce Diatom Classification

📅 2026-03-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of visual classification in decentralized settings with scarce labels and heterogeneous class distributions, such as diatom taxonomy, where clients exhibit only partial overlap in label spaces. To tackle this, the authors propose a self-supervised federated learning framework featuring two key components: PreDi, a controlled partitioning scheme that decouples label heterogeneity into two orthogonal dimensions—class popularity and local label set size—and PreP-WFL, a personalized weighted aggregation strategy that enhances representations of low-popularity classes during both contrastive pretraining and federated fine-tuning. Experimental results demonstrate that the proposed method consistently outperforms local training under both homogeneous and heterogeneous settings, substantially mitigating performance degradation caused by label space misalignment, with particularly pronounced gains on rare classes.

Technology Category

Machine Learning: Distributed Machine Learning & Federated LearningSearch and Optimization: Distributed SearchComputer Vision: Representation Learning for Vision

Application Category

User Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSystems and Infrastructure for Web, Mobile and WoT: Decentralized Web and Fediverse systems
📝 Abstract
Label-scarce visual classification under decentralized and heterogeneous data is a fundamental challenge in pattern recognition, especially when sites exhibit partially overlapping class sets. While self-supervised federated learning (SSFL) offers a promising solution, existing studies commonly assume the same data heterogeneity pattern throughout pre-training and fine-tuning. Moreover, current partitioning schemes often fail to generate pure partially class-disjoint data settings, limiting controllable simulation of real-world label-space heterogeneity. In this work, we introduce SSFL for diatom classification as a representative real-world instance and systematically investigate stage-specific data heterogeneity. We study cross-site variation in unlabeled data volume during pre-training and label-space misalignment during downstream fine-tuning. To study the latter in a controllable setting, we propose PreDi, a partitioning scheme that disentangles label-space heterogeneity into two orthogonal dimensions, namely class Prevalence and class-set size Disparity, enabling separate analysis of their effects. Guided by the resulting insights, we further propose PreP-WFL (Prevalence-based Personalized Weighted Federated Learning) to adaptively strengthen rare-class representations in low-prevalence scenarios. Extensive experiments show that SSFL consistently outperforms local-only training under both homogeneous and heterogeneous settings. The pronounced heterogeneity in unlabeled data volume is associated with improved representation pre-training, whereas under label-space heterogeneity, prevalence dominates performance and disparity has a smaller effect. PreP-WFL effectively mitigates this degradation, with gains increasing as prevalence decreases. These findings provide a mechanistic basis for characterizing label-space heterogeneity in decentralized recognition systems.
Problem

Research questions and friction points this paper is trying to address.

self-supervised federated learning
data heterogeneity
label-scarce classification
diatom recognition
label-space heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Federated Learning
Label-Space Heterogeneity
PreDi Partitioning
Prevalence-based Personalization
Diatom Classification
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Mingkun Tan
Biodata Mining Group, Faculty of Technology, University of Bielefeld, Germany
X
Xilu Wang
Computer Science Research Centre, University of Surrey, United Kingdom
M
Michael Kloster
Phycology Group, Faculty of Biology, University of Duisburg-Essen, Germany
Tim W. Nattkemper
Tim W. Nattkemper
Professor of Computer Science, Bielefeld University
BioinformaticsBioimage InformaticsUnderwater Image AnalysisVisualizationData Mining