🤖 AI Summary
This study addresses the challenge of intuitively inspecting and intervening in task parsing and coordination within large language model-driven multi-robot systems by developing a mixed-reality control framework. Methodologically, it designs a four-stage supervision workflow under natural language instructions and introduces a novel multi-scale intervention mechanism based on structured commitments. By integrating distributed state synchronization algorithms, the system enables coordinated presentation of local and global views to enhance transparency. Experimental results demonstrate that the proposed framework significantly reduces user cognitive load while effectively improving situational awareness, human-robot trust, and the perceived controllability over multi-robot teams.
📝 Abstract
Large language models (LLMs) let users direct heterogeneous multi-robot systems (MRS) through natural language, but make task interpretation, robot assignment, and coordination difficult to inspect and change. Based on a formative study with 12 non-expert users, we developed MRPilot, a mixed reality system organized around four stages of supervision and intervention. MRPilot represents robot-team plans and execution states as structured commitments shared across synchronized situated and overview views. Across four stages, it helps users resolve ambiguous references (Forming), review plans before execution (Reviewing), monitor distributed execution (Following), and make robot-level or team-level changes when problems arise (Repairing). In a within-subjects study with 20 participants in a virtual reality-simulated home, MRPilot reduced workload, increased situational awareness, transparency, trust, and perceived control compared with a conventional LLM-based conversational interface using the same LLM planner and robot capabilities. We provide design implications for multi-scale intervention, adaptive supervision, and calibrated reliance in LLM-based MRS.