PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation

📅 2026-04-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the issue of identity privacy leakage in centralized training for personalized talking-head generation by proposing a privacy-preserving framework based on federated learning. Each client trains a lightweight LoRA identity adapter locally using private audiovisual data, while sharing a common diffusion backbone model without uploading raw data. The approach introduces an identity-stable federated aggregation mechanism and temporal denoising consistency regularization to effectively mitigate inter-frame flickering and identity drift. By integrating secure aggregation with client-side differential privacy, the method achieves high-quality, temporally coherent personalized synthesis while rigorously protecting user privacy. Experimental results demonstrate the feasibility and superiority of the proposed framework in resource-constrained settings.

Technology Category

Machine Learning: PrivacyComputer Vision: Diffusion Models for VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

User Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSecurity and Privacy: Privacy-enhancing technologiesResponsible Web: Data and user privacy-enhancing technologies for the Web
📝 Abstract
Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech datasets, raising major privacy concerns. The problem is more acute for personalized talking-head generation, where identity-specific data are highly sensitive and often cannot be pooled across users or devices. PrivFedTalk is presented as a privacy-aware federated framework for personalized talking-head generation that combines conditional latent diffusion with parameter-efficient identity adaptation. A shared diffusion backbone is trained across clients, while each client learns lightweight LoRA identity adapters from local private audio-visual data, avoiding raw data sharing and reducing communication cost. To address heterogeneous client distributions, Identity-Stable Federated Aggregation (ISFA) weights client updates using privacy-safe scalar reliability signals computed from on-device identity consistency and temporal stability estimates. Temporal-Denoising Consistency (TDC) regularization is introduced to reduce inter-frame drift, flicker, and identity drift during federated denoising. To limit update-side privacy risk, secure aggregation and client-level differential privacy are applied to adapter updates. The implementation supports both low-memory GPU execution and multi-GPU client-parallel training on heterogeneous shared hardware. Comparative experiments on the present setup across multiple training and aggregation conditions with PrivFedTalk, FedAvg, and FedProx show stable federated optimization and successful end-to-end training and evaluation under constrained resources. The results support the feasibility of privacy-aware personalized talking-head training in federated environments, while suggesting that stronger component-wise, privacy-utility, and qualitative claims need further standardized evaluation.
Problem

Research questions and friction points this paper is trying to address.

talking-head generation
privacy
personalization
federated learning
identity-sensitive data
Innovation

Methods, ideas, or system contributions that make the work stand out.

federated learning
diffusion models
identity-stable adapters
privacy-preserving generation
temporal consistency
🔎 Similar Papers
No similar papers found.
Soumya Mazumdar
Soumya Mazumdar
University of Canberra
Public health & geographyBuilt EnvironmentRefugee health equityGISGreenspace
V
Vineet Kumar Rakesh
Engineering Sciences, Homi Bhabha National Institute, Anushaktinagar, Mumbai, Maharashtra 400094, India; Computer and Informatics Group / Radioactive Ion Beam Facilities Group, Variable Energy Cyclotron Centre, 1/AF, Bidhannagar, Kolkata 700064, West Bengal, India
T
Tapas Samanta
Engineering Sciences, Homi Bhabha National Institute, Anushaktinagar, Mumbai, Maharashtra 400094, India; Computer and Informatics Group / Radioactive Ion Beam Facilities Group, Variable Energy Cyclotron Centre, 1/AF, Bidhannagar, Kolkata 700064, West Bengal, India