🤖 AI Summary
This study addresses the limitations of passive execution paradigms in embodied intelligence under partially observable scenarios by proposing an active exploration framework based on multimodal large models. Methodologically, the framework leverages large models for high-level task reasoning and integrates visual perception for scene understanding. It further introduces a fine-grained perception-action interleaving strategy, establishing a synergistic closed loop among planning, perception, and control modules to enable robots to dynamically interact with the environment and adaptively acquire information. Validation on real-world "find-and-place" tasks demonstrates that the proposed framework effectively overcomes challenges such as initially invisible targets and environmental distractors, significantly enhancing the robustness of autonomous exploration and manipulation capabilities in complex, unknown environments.
📝 Abstract
Recent advances in agentic systems have substantially enhanced the long-horizon capability of embodied manipulation. However, many existing frameworks still follow a passive execution paradigm, which limits their applicability to real-world scenarios involving textual semantic cues, distractors, and initially invisible targets. To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions. Specifically, our framework consists of three collaborative modules: a planning module for high-level task reasoning, a perception module for visual scene understanding, and an execution module for low-level manipulation. This design allows the robot to actively acquire task-relevant information, adapt its behavior based on environmental feedback, and complete manipulation tasks under partial observability. Furthermore, we introduce a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to improve exploration robustness. We evaluate our method on a realistic Find-and-Place task, demonstrating its effectiveness in challenging environments where target objects must be actively discovered before manipulation.