LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决机器人操作中部分可观测性问题,提出LIFD框架,通过单目RGB视图和循环记忆学习并补全3D场景表示,提高操作成功率。
📝 Abstract
Robotic manipulation under partial observability requires spatial information that extends beyond the current view. Geometry-aware RGB features describe visible structure, but previously observed regions may disappear as the robot or scene moves. Maintaining a useful scene representation therefore requires retaining observation history while inferring missing content without losing its connection to visible evidence. We introduce LIFD (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory. LIFD learns a scene-token representation from multi-view agreement and completes it from a single RGB view and recurrent memory. A rectified-flow model generates the tokens while Anchor-Guided Cross-Attention conditions completion on current geometric features. Compact slot features connect this representation to a manipulation policy. Multi-view and geometric supervision are used during representation learning; deployment requires one RGB camera, proprioception, and a task instruction. LIFD (Staged) reaches 91.6% average success on LIBERO and 79.8% on MetaWorld, improving LIBERO average success by 3.1 percentage points over Joint training. On four UR5e task families with ten demonstrations per family, it achieves 56.0% mean success, compared with 40.5% for OpenVLA-7B.
Problem

Research questions and friction points this paper is trying to address.

partial observability
spatial information
geometry-aware RGB features
scene representation
observation history
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D-aware Scene Memory
Anchor-Guided Cross-Attention
Multi-view Agreement
Scene-token Representation
Rectified-flow Model
🔎 Similar Papers
No similar papers found.