Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the conflict between representation drift and parameter aggregation in multi-robot federated learning under non-IID data distributions by proposing the FedDRMan framework. To mitigate representational shifts caused by heterogeneous data, the method introduces low-rank multimodal subspace distillation. It further enhances collaborative efficiency through a client-compatible grouped clustering strategy and ensures global model convergence stability via a spectral rebalancing aggregation mechanism. By integrating federated learning, behavior cloning, knowledge distillation, and vision-language-action models, this work achieves a peak success rate of 80.7% on the LIBERO benchmark, outperforming the strongest baseline by 11.6 percentage points. These results demonstrate that FedDRMan effectively resolves the challenge of collaborative policy learning for multi-robot systems operating with heterogeneous data.
📝 Abstract
Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter aggregation destructive. We present FedDRMan, a federated subspace-guided distillation framework for heterogeneous robot manipulation. At each communication round, the server model provides a frozen teacher for local behavior cloning, while low-rank multimodal subspace and action-distribution distillation preserve globally useful representation geometry and policy behavior. To address heterogeneous aggregation, FedDRMan groups clients by update compatibility and maintains a persistent model for each cluster. The server then spectrally rebalances each compatible aggregate to mitigate attenuation of weaker task-relevant robot-policy update directions. Extensive experiments on LIBERO across diverse non-IID settings, heterogeneity levels, client participation variation, together with ablations and aggregation analyses, show that FedDRMan substantially improves knowledge transfer and consistently outperforms strong federated baselines achieving a peak mean success rate of 80.7%, 11.6 percentage points above the strongest evaluated federated baseline.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Non-IID
Multi-Robot Manipulation
Representation Drift
Heterogeneous Aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
Policy Distillation
Vision-Language-Action
Non-IID
Multi-Robot Manipulation