🤖 AI Summary
This study addresses the problem of action error accumulation and out-of-distribution states in generative robot policies caused by heterogeneous demonstrations. To this end, we propose TeV, a temporal verification framework built upon a flow-matching vision-language-action model. By introducing time-aware tokens to compress historical observations and actions, TeV enables trajectory-level continuity assessment. Furthermore, it constructs positive and negative samples without expert annotations via energy-based contrastive learning, leveraging the resulting energy landscape to guide intermediate sampling optimization. Experiments demonstrate that TeV significantly improves task success rates in both simulated and real-world environments, generates smoother execution trajectories, and achieves reliable ranking of candidate actions.
📝 Abstract
Generative policies have emerged as a promising paradigm for robot learning, combining expressive generative action modeling with scalable imitation learning from large demonstration corpora. However, heterogeneous demonstrations can induce suboptimal action chunks whose errors compound over time, eventually driving the robot into out-of-distribution states from which recovery is difficult. Action verification offers a test-time scaling strategy for mitigating this failure mode by sampling multiple candidate actions and using a verifier to select one for execution. Existing approaches, however, remain temporally myopic and costly to train, evaluating candidates from the current observation alone without accounting for trajectory continuity and often relying on large verifiers and additional expert demonstrations. In this paper, we introduce Temporal Verification (TeV), an efficient temporally aware action verification framework for flow-matching VLAs. TeV first learns a temporal token that summarizes recent observation--action history, enabling candidate chunks to be evaluated as continuations of the execution trajectory rather than as isolated predictions. Conditioned on this token, TeV constructs positive--negative pairs without additional expert demonstrations or preference annotations and trains an energy-based verifier contrastively to assign lower energy to higher-quality, trajectory-consistent action chunks. Beyond post-hoc ranking, TeV further uses the learned energy landscape to guide intermediate flow samples toward lower-energy regions, improving candidates before final selection. Extensive experiments in simulation and real-world settings demonstrate that TeV provides reliably ranks action candidates, improves task success rates, and produces smoother execution trajectories.