🤖 AI Summary
This study addresses the ambiguity, logical conflicts, and redundant execution caused by existing replanning methods when domestic robots receive new instructions during task execution. We propose CIRRA, a novel framework that introduces a dual-level architecture integrating large language model-based semantic reasoning with rule-constrained structural organization to achieve continuous instruction coordination. This mechanism preserves the execution backbone while evaluating only modified segments, thereby maintaining contextual alignment and minimizing redundancy. Experimental results demonstrate that CIRRA achieves a decision consistency rate of 74.2% on the CHIRP benchmark, surpassing baselines by over 30%. Furthermore, real-world robot experiments yield zero erroneous interruptions, comprehensively outperforming existing approaches.
📝 Abstract
Household robots must accommodate new user instructions while executing ongoing tasks. Existing agents often regenerate or extensively revise the remaining task sequence, introducing plan ambiguity, logical inconsistency, and redundant execution. We formulate continual instruction reconciliation and propose CIRRA (Continual Instruction Reconciliation for Robot Agents), a dual-level framework combining LLM-based semantic reasoning with rule-constrained structural integration. CIRRA first grounds incoming instructions to unique executable skills and resolves underspecified actions and execution locations. It then preserves the ongoing subtask sequence as an execution backbone and generates integration candidates by inserting incoming subtasks into location-matched segments. The semantic reasoner evaluates only modified segments to identify dependencies and conflicts and select the most logically coherent local integration. This structure-preserving process maintains alignment with ongoing execution, mitigates ambiguity and inconsistency, and reuses shared subtasks to reduce redundant execution. We also introduce CHIRP (Continual Household Instruction Reconciliation and Planning), a text-based benchmark of 120 episodes across eight household environments and six categories of everyday activities. On CHIRP, CIRRA achieves 74.2% decision agreement, exceeding the strongest replanning baseline by 30 percentage points; every correct fusion decision yields a correctly placed, conflict-free schedule. On a Unitree G1 humanoid, CIRRA interrupts ongoing skills at the correct moment in every trial and significantly outperforms all baselines on every metric.