RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing robotic datasets—namely high acquisition costs, restricted morphological diversity, and insufficient fine-grained structured annotations needed for generalization and long-horizon dynamics modeling—by introducing a unified intermediate representation for robotic manipulation and embodied world modeling. This representation serves, for the first time, as a bidirectional interface: it standardizes the low-level action space while also constraining the simulation dynamics of open-world physics engines. Leveraging this framework, the authors construct a large-scale dataset comprising over 230,000 manipulation episodes across 571 scenes, featuring dense frame-level annotations, spatiotemporal embodied visual-language question answering, integrated VLM/VLA models, and controllable conditional future prediction. Experiments demonstrate that the proposed approach substantially enhances model performance and generalization in intermediate representation reasoning, task execution, and world dynamics modeling.
📝 Abstract
Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, execution, or long-horizon environment dynamics simulation. Building on our prior work, RoboInter1.0, we present RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling. RoboInter1.5 provides a unified resource of data, benchmarks, and models centered on dense manipulation-oriented intermediate representations. Specifically, RoboInter-Data contains over 230k manipulation episodes across 571 scenes with dense per-frame annotations covering more than ten types of intermediate representations, including subtasks, primitive skills, object and gripper grounding, segmentation, affordance, grasp poses, contact points, motion traces, etc. Built upon these annotations, RoboInter-VQA introduces spatial and temporal embodied VQA tasks to benchmark and improve the intermediate-representation reasoning capabilities of our RoboInter-VLM. RoboInter-VLA further studies how such representations benefit action execution through implicit, explicit, and modular plan-then-execute paradigms. To better model the physical world, we further introduce RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states. Extensive evaluations demonstrate that RoboInter1.5 provides a unified spatiotemporal scaffolding for intermediate representations. Rather than treating intermediate representations merely as interpretable signals, RoboInter1.5 conceptualizes them as a bidirectional interface that both regularizes low-level action spaces and constrains the latent rollouts of open-world physical simulators.
Problem

Research questions and friction points this paper is trying to address.

robot datasets
embodied world modeling
intermediate representations
fine-grained annotation
generalizable reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

intermediate representation
embodied reasoning
robotic manipulation
world modeling
dense annotation