From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of physical skill reuse and the coupling between high- and low-level control in multi-agent collaboration by proposing a hierarchical reinforcement learning framework. Through object-centric action space distillation, single-agent policies are transformed into reusable object manipulation skills. By freezing the low-level actuators and generating region-based proxy actions, the proposed method effectively decouples contact execution from high-level coordination. This approach facilitates composable policy learning across diverse geometries, tasks, and team sizes. Experimental evaluations on various interaction tasks demonstrate the robustness of the acquired skills, yielding substantial improvements in both policy reusability and collaborative efficiency.
📝 Abstract
Physics-based human-object interaction has achieved robust single-agent manipulation skills, yet extending them to multi-agent cooperative tasks remains challenging. Existing approaches typically adapt interaction policies through task-specific fine-tuning, which entangles low-level contact-rich execution with high-level coordination and limits reuse across object geometries, interaction types, and team sizes. We propose a hierarchical framework that converts a single-agent HOI policy into a reusable Object-oriented Motion Skill. Specifically, we reinterpret teacher rollouts as object-oriented action supervision by extracting short-horizon object-proxy motions from executed trajectories, and distill task-specific teachers into a low-level skill operating in an Object-oriented Action Space. For downstream tasks, the distilled skill is frozen as a reusable executor, while a high-level policy coordinates multiple agents by generating region-wise object-oriented actions conditioned on the shared object, task goal, agent states, and local manipulation regions. This formulation shifts multi-agent HOI learning from direct contact-rich full-body control to compact object-level proxy-motion coordination. Experiments on diverse HOI tasks show that the distilled Object-oriented Motion Skill supports robust proxy-motion execution and enables composable policy learning across different interaction types, object geometries, and team sizes.
Problem

Research questions and friction points this paper is trying to address.

Human-Object Interaction
Multi-Agent Cooperation
Policy Reusability
Hierarchical Framework
Composable Policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Framework
Object-oriented Motion Skill
Multi-Agent Human-Object Interaction
Policy Distillation
Composable Policy Learning
🔎 Similar Papers
No similar papers found.