Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of domain heterogeneity hindering personalized fine-tuning in federated vision-language models by proposing MOSAIC. Departing from the conventional class-agnostic domain shift assumption, this work pioneers the modeling of fine-grained semantic-appearance interactions and quantifies their detrimental effects. Specifically, MOSAIC effectively mitigates the degradation of decision boundaries caused by class-domain interactions through a decision-aware harmfulness score and low-rank residual adapters. Technically, it integrates parameter-efficient fine-tuning, low-rank decomposition, image-conditioned gating, and harmful-pair-aware reweighting strategies. Extensive experiments on benchmarks such as Office31 demonstrate that MOSAIC significantly improves macro client-level Top-1 accuracy, consistently outperforming existing baselines.
📝 Abstract
Federated parameter-efficient fine-tuning enables distributed clients to adapt pretrained vision-language models without sharing raw data or updating the full backbone. Its effectiveness, however, is limited by domain heterogeneity across clients. Existing personalized methods separate globally shared knowledge from client-specific style, but they largely treat each domain as a class-agnostic transformation. We show that this abstraction is insufficient: the cross-domain displacement associated with a fixed domain varies across semantic classes, and only a subset of these class-domain residuals damages the image-text decision margin. We therefore propose Margin-Oriented Semantic-Appearance Interaction Correction (MOSAIC), which first constructs a decision-aware harmfulness score that measures whether a training-derived class-domain residual favors a competing text prototype over the true class. It then models fine-grained class-domain interactions with a low-rank residual adapter whose class factors and residual basis are globally shared while domain factors remain client-private. An image-conditioned gate further controls candidate-wise correction, and harmful-pair-aware reweighting prioritizes decision-relevant residuals during local optimization. Extensive experiments on Office31, OfficeHome, and DomainNet100 demonstrate that MOSAIC consistently improves macro-client top-1 accuracy across all evaluated domain-shift and joint domain-label-shift settings.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Vision-Language Models
Domain Heterogeneity
Personalized Federated Learning
Decision Margin
Innovation

Methods, ideas, or system contributions that make the work stand out.

Personalized Federated Learning
Vision-Language Models
Parameter-Efficient Fine-Tuning
Low-Rank Residual Adapter
Decision Margin Correction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wentao Yue
Q
Qingyu Mao
T
Tianyou Lai
A
Ahmed M. Abdelmoniem
Qilei Li
Qilei Li
Central China Normal University
Deep learningComputer science