🤖 AI Summary
该研究提出了一种闭环框架,通过整合视觉语言模型观测、抓取可行性和有限视界信念空间规划来解决部分可观测条件下基于语言指导的对象检索问题。
📝 Abstract
Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability.