MAAP: Multi-Agent Active Perception for Collaborative Manipulation

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过MAAP方法,让每个机械臂在执行操作的同时利用腕部摄像头提供移动视角,并结合RAIL控制器预测角色和动作,提高多任务协作成功率。
📝 Abstract
Multi-agent manipulation naturally produces multiple task-driven viewpoints: every arm carries a wrist camera and moves through the scene while acting. Yet these observations are typically underutilized, and active perception in manipulation is still often treated as requiring a dedicated sensing agent. We introduce MAAP (Multi-Agent Active Perception), in which every arm is dual-purpose: it executes manipulation actions and, through the wrist camera it carries, simultaneously serves as a moving viewpoint for the team. We pair this with RAIL (Role-Aware Imitation Learning), a controller that predicts each arm's current role alongside its action chunk and conditions action generation on it, representing role-dependent actions within one network. Across four simulated tasks, widening the perception regime lifts average success from 56.5% with a fixed camera to 62.5% with one active wrist view and 70.0% with all of them, while MAAP+RAIL reaches 79.2%. RAIL's additional gain is concentrated on the three-arm Microwave task, where success rises from 47% to 82% on identical multi-wrist inputs. On a dual-arm platform, MAAP+RAIL succeeds in 14 of 20 placement trials compared with 0 of 20 for fixed-view ACT. Collaborative manipulation can thus serve as an active perception mechanism in its own right.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Active Perception
Collaborative Manipulation
Wrist Camera
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Active Perception
Role-Aware Imitation Learning
collaborative manipulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bruno N. Y. Chen
Carnegie Mellon University
Li Kang
Li Kang
Shanghai Jiao Tong University
Embodied AIMulti-Agent System
Heng Zhou
Heng Zhou
Jiangnan University
Multi-modal LearningImage ProcessingComputer VisionRemote Sensing
Xiufeng Song
Xiufeng Song
Shanghai Jiao Tong University
Computer VisionEmbodied Intelligence
Z
Zhemeng Zhang
The University of Hong Kong
J
Jiahua Ma
Sun Yat-sen University
Y
Yiran Qin
The Chinese University of Hong Kong, Shenzhen