Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unreliability of existing robotic manipulation policies, where the absence of explicit outcome prediction mechanisms leads to long-horizon error accumulation. We propose a "simulate-before-act" paradigm by constructing a geometry-aware spatiotemporal action representation framework. Specifically, our approach integrates 3D Gaussian Splatting, SAM3D, and vision-language models to generate high-fidelity simulated environments. By coupling spatiotemporal reasoning with trajectory planning, the system simulates and evaluates potential actions prior to execution, filtering out infeasible behaviors and enabling geometry-based terminal state estimation. Extensive experiments demonstrate that this method significantly improves manipulation reliability across rigid, soft, and deformable objects. These findings validate the scalability of combining geometric reconstruction with high-quality simulation for robust robotic control.
📝 Abstract
Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desired final states, causing execution errors to accumulate during long-horizon manipulation. We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment. Our approach constructs a high-fidelity robotic environment from RGB-D observations using 3D Gaussian Splatting and SAM3D, enabling ``simulation and evaluation before acting''. It integrates visual observations and language instructions with spatio-temporal reasoning for long-horizon task planning using large vision-language models. To bridge high-level planning and real-world execution, we introduce Gaussian-aware final-state estimation through geometric sampling and state-based trajectory planning. Before execution, candidate action sequences are simulated and evaluated in the Gaussian-SAM environment to filter infeasible behaviours. We validate our approach on representative manipulation tasks involving rigid, soft, and deformable objects, including cube placing, toy packing, and duck rearrangement, demonstrating that geometry-aware spatio-temporal reasoning and state-aware execution improve manipulation reliability across different object categories. Our results suggest that combining geometry-aware reconstruction with high-quality rendering and simulation provides a scalable approach for evaluating robotic manipulation behaviours. Website: https://robot-gst.github.io
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
policy evaluation
error accumulation
long-horizon tasks
task outcome prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Gaussian Splatting
Spatio-temporal reasoning
Real-to-sim verification
Vision-language models
Final-state estimation
💼 Related Jobs
No related jobs found.
S
Sichao Liu
KTH
Z
Zekun Wang
KTH
L
Lixuan Tang
EPFL
Yiming Li
Yiming Li
Idiap Research Institute, EPFL
machine learningrobot perceptionmotion planningmanipulation
X
Xiaohan Wang
Beihang University
H
Hanzhi Zhang
KTH
D
Daqiang Guo
The Hong Kong University of Science and Technology (Guangzhou)
P
Peng Zhou
Great Bay University
Lihui Wang
Lihui Wang
Chair Professor of Sustainable Manufacturing, KTH
AI in manufacturinghuman-robot collaborationsmart manufacturing systems