GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of Vision-Language-Action (VLA) models in unseen tasks, novel configurations, and long-horizon scenarios by proposing a steerable framework. The approach decouples semantic acquisition, trajectory generation, and action execution. It leverages the commonsense priors of general-purpose vision-language models to synthesize 2D visual trajectories, which subsequently condition low-level action policies, thereby enabling effective transfer from high-level semantics to low-level control. Additionally, a Mixture-of-Experts architecture is incorporated to enhance system robustness. Experiments conducted on the LIBERO benchmark and physical robot platforms demonstrate that the proposed method significantly improves cross-scenario generalization performance compared to existing VLA baselines.
📝 Abstract
Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to follow novel language instructions. Off-the-shelf vision-language models (VLMs) often provide stronger generalization, but cannot directly control robot actions. To combine the common sense of VLMs with VLA control, we propose Guided Trace VLA (GT-VLA), a steerable framework that accepts guidance from an external generalist VLM through trace-conditioned action generation. GT-VLA uses a generalist model to identify semantic guidance for the current skill, converts this guidance into a 2D visual trace, and conditions its action policy on the resulting trace-rendered observation. This design separates semantic target acquisition, trace generation, and low-level action execution, allowing high-level guidance to propagate to robot actions. GT-VLA uses a Mixture-of-Experts architecture with skill-specific trace and action modules for robust execution. We evaluate GT-VLA on LIBERO and a physical robot platform, showing improved generalization over recent VLA baselines in both settings. The code and additional supplemental materials are available on our project website at https://ivaniz.github.io/gt-vla/.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
generalization
robotic manipulation
vision-language models
unseen tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Trace Guidance
Mixture-of-Experts (MoE)
Robotic Manipulation
Generalization