🤖 AI Summary
This study addresses the unreliability and computational inefficiency of large reasoning models in robotic planning, where valid plans are frequently overwritten or constraint violations persist. To mitigate these issues, this work proposes a verifier-guided framework that introduces a novel mechanism for exposing and validating intermediate plans without interrupting decoding trajectories. Specifically, an inference-time monitor is constructed to preserve valid intermediate plans during generation and guide iterative error correction. The proposed approach significantly improves planning success rates, accelerates error rectification, and reduces token consumption. The effectiveness of this method has been validated through experiments conducted on both simulated and real-world robotic manipulators.
📝 Abstract
Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.