Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient physical dynamics priors and the difficulty of integrating semantics with dynamics in Vision-Language-Action (VLA) models by proposing ACT3, a three-stream Transformer architecture. The method introduces an innovative design featuring independent context streams coupled with a shared action interaction mechanism. Through hierarchical attention, semantic representations from vision-language models and dynamic features from world models are injected into action experts, enabling effective complementarity between semantic understanding and dynamic prediction. Experimental results demonstrate that ACT3 significantly outperforms existing methods on both simulated and real-world robotic benchmarks, offering an efficient architectural paradigm for embodied intelligent control.
📝 Abstract
Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs). However, such VLM backbones offer insufficient physical dynamics priors, which limits the generalization capabilities of robot policies. Recent efforts therefore integrate video-generation World Models (WMs) into robot policies through various strategies, using predictive dynamics to facilitate action generation. Despite these advances, harnessing semantic understanding and dynamics prediction as complementary guidance for action generation remains challenging. In this paper, we introduce $\mathrm{ACT}^3$, a simple yet effective Action-Centric Tri-Stream Transformer that fuses semantic and dynamics information into control actions while preserving the distinct roles of context streams. Specifically, $\mathrm{ACT}^3$ enables the dedicated action expert to access VLM and WM representations through layerwise attention, with each backbone attending only within its own stream. This straightforward interaction design maintains independent forward propagation in the context streams while allowing both backbones to be updated through control supervision. Experiments on both simulated and real-world robotic manipulation benchmarks show that the proposed $\mathrm{ACT}^3$ yields results superior to its counterparts.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
physical dynamics priors
world models
robotic manipulation
action generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Model
Tri-Stream Transformer
World Model
Layerwise Attention
Robotic Manipulation
🔎 Similar Papers
No similar papers found.