Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of parsing and reusing temporal task structures in video-based policy learning by proposing an interpretable framework grounded in Signal Temporal Logic (STL). Methodologically, a vision-language model extracts semantic event trajectories from observation-only videos to construct parameterized STL specifications, which are symbolized via robot trajectory-calibrated thresholds. A novel short- and long-term temporal decoupling mechanism synergistically optimizes policies using dense short-term rewards alongside long-term causal progress-monitoring rewards, facilitating cross-embodiment transfer. Experimental results demonstrate an average success rate of 85.8% on manipulation tasks, while quadruped robots achieve 100% success across multiple velocities with superior energy efficiency.
📝 Abstract
Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves $85.8\%$ average success-once and $67.0\%$ success-at-end, compared with $81.5\%/59.5\%$ for native dense PPO and $65.0\%/42.3\%$ for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve $100\%$ success across velocities from $0.3$ to $2.1\,\mathrm{m/s}$ while remaining competitive in high-speed energy efficiency. Project webpage: \href{https://video2stl.github.io/}{video2stl}.
Problem

Research questions and friction points this paper is trying to address.

video-based policy learning
temporal specifications
robot learning
cross-embodiment transfer
Signal Temporal Logic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Signal Temporal Logic
Vision-Language Model
Cross-embodiment Transfer
Policy Learning
Temporal Specifications
Merve Atasever
Merve Atasever
University of Southern California
Machine LearningReinforcement LearningRoboticsDifferential Geometry
K
Keyan Azbijari
Department of Computer Science, University of Southern California
C
Cagan Bakirci
Department of Computer Science, University of Southern California
B
Bo-Ruei Huang
Department of Computer Science, University of Southern California
T
Tolga Izdas
Department of Computer Science, University of Southern California
Z
Zahra Shahrooei
Department of Computer Science, University of Southern California
R
Richard Yang
University of Florida
E
Erdem Biyik
Department of Computer Science, University of Southern California
Jyotirmoy V. Deshmukh
Jyotirmoy V. Deshmukh
Associate Professor, University of Southern California
Cyberphysical systemsFormal/Statistical VerificationTemporal logicAI safetyReinforcement learning