Directed Temporal Representations for Offline Visual Control

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation in offline visual control where latent geometry focuses solely on predictive similarity while lacking temporal reachability alignment. To overcome this, it proposes constructing a directed temporal quasi-metric over frozen LeWM features to estimate arrival costs. This work leverages temporal progress signals as critic objectives to directly optimize goal-conditioned policies. By integrating short-horizon calibration, long-horizon bootstrapping, and action consistency constraints, the method enables direct policy learning without iterative search at test time. Experiments across ten visual control tasks demonstrate superior performance, effectively validating that temporal supervision enhances policy optimization while preserving structural consistency.
📝 Abstract
Predictive world models provide compact visual representations for control. Control requires a latent geometry aligned with temporal reachability rather than predictive similarity alone. We introduce Directed Temporal Representations for Control (DTRC), which learns such a geometry from offline visual trajectories on top of frozen LeWorldModel (LeWM) features. DTRC constructs a directed temporal quasimetric over the learned control representation. Short-range temporal offsets calibrate the distance scale. Bootstrapped targets extend temporal reachability across longer horizons. Action-conditioned consistency aligns the representation with local transition dynamics. The resulting distance estimates temporal reaching cost, and its change across a transition defines goal-relative temporal progress. We use this progress signal as a temporal critic for direct goal-conditioned policy learning. Model-assisted targets provide an additional training-time refinement under behavior-support and dynamics-agreement constraints. Across ten visual control tasks, DTRC achieves strong goal-conditioned control performance relative to planning and direct-policy baselines. Held-out diagnostics on the four LeWM tasks show consistent short-range temporal calibration, task-dependent long-range and directional structure, and positive transition-level progress. Temporal supervision improves the same flow-policy parameterization across all four LeWM tasks, while the resulting policy acts directly without iterative trajectory search at test time.
Problem

Research questions and friction points this paper is trying to address.

offline visual control
temporal reachability
world models
goal-conditioned policy
latent geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

directed temporal quasimetric
offline visual control
goal-conditioned policy learning
temporal progress critic
world model representations
🔎 Similar Papers
No similar papers found.