From Dialogue to Execution: Mixture-of-Agents Assisted Interactive Planning for Behavior Tree-Based Long-Horizon Robot Execution

📅 2026-03-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a novel framework integrating Mixture-of-Agents (MoA) with Behavior Trees to address the limitations of existing large language model–based interactive planning in long-horizon robotic tasks, which often rely on frequent human-in-the-loop queries and produce flat, inflexible plans ill-suited for complex logic. By introducing MoA into robotic task planning for the first time, the approach leverages multiple expert agents to collaboratively generate hierarchical task structures, while Behavior Trees enable dynamic policy switching and autonomous failure recovery. Experiments on a bartending task demonstrate a 27% reduction in human intervention, with generated behavior trees exhibiting high structural and semantic fidelity to fully handcrafted counterparts. The method’s robustness and scalability are further validated by successful execution of extended beverage preparation sequences on a physical robot.

Technology Category

Humans and AI: Human-Aware Planning and Behavior PredictionMultiagent Systems: Multiagent PlanningIntelligent Robots: Behavior Learning & Control

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsResponsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Agentic search
📝 Abstract
Interactive task planning with large language models (LLMs) enables robots to generate high-level action plans from natural language instructions. However, in long-horizon tasks, such approaches often require many questions, increasing user burden. Moreover, flat plan representations become difficult to manage as task complexity grows. We propose a framework that integrates Mixture-of-Agents (MoA)-based proxy answering into interactive planning and generates Behavior Trees (BTs) for structured long-term execution. The MoA consists of multiple LLM-based expert agents that answer general or domain-specific questions when possible, reducing unnecessary human intervention. The resulting BT hierarchically represents task logic and enables retry mechanisms and dynamic switching among multiple robot policies. Experiments on cocktail-making tasks show that the proposed method reduces human response requirements by approximately 27% while maintaining structural and semantic similarity to fully human-answered BTs. Real-robot experiments on a smoothie-making task further demonstrate successful long-horizon execution with adaptive policy switching and recovery from action failures. These results indicate that MoA-assisted interactive planning improves dialogue efficiency while preserving execution quality in real-world robotic tasks.
Problem

Research questions and friction points this paper is trying to address.

interactive planning
long-horizon tasks
human-in-the-loop
task representation
robot execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Agents
Behavior Trees
Interactive Planning
Long-Horizon Execution
LLM-based Robotics