Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

📅 2026-09-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出时间-频率几何交叉注意力(TFGCA)模块,解决视觉-语言-动作模型中动作序列的频率和跨阶段几何结构问题,提升模型性能。
📝 Abstract
Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is essentially a short multivariate trajectory, but inside these models it is a sequence of generic per-timestep hidden tokens decoded by a linear head. This under-serves two motion structures. First, frequency: a chunk superimposes a smooth global trend and fine corrective motion across time scales, and a single token entangles them. Second, cross-phase geometry: motions of different phases (reach, contact, grasp adjustment, settling) unfold along very different, near-orthogonal directions in representation space, yet are tightly related for the task and arise across the time axis. Dot-product attention scores alignment by an inner product, so it favors aligned tokens and is least sensitive near orthogonality, leaving such relationships for the network to recover through a detour. We introduce Time-Frequency Geometric Cross-Attention (TFGCA), a drop-in module repairing both blind spots. TFGCA uses a per-dimension learnable stationary wavelet transform to decompose the action chunk into time-frequency tokens, and each time token retrieves information from them via a cross-attention that fuses the dot product (similarity) with the wedge-product magnitude (sensitive to near-orthogonality) through a learnable weight. A zero-initialized residual reproduces the base behavior at initialization, so it can be dropped onto a pretrained VLA and fine-tuned jointly. Relative to the same-source base, TFGCA improves in-distribution LIBERO by +1.5 on average, the OOD LIBERO-Plus by +6.3, the randomized average under RoboTwin domain randomization by +28.5, and the overall success rate on three real-robot AgiBot A2 tasks by +11.67 points, with larger gains out of distribution.
Problem

Research questions and friction points this paper is trying to address.

Time-Frequency
Cross-Phase Geometry
Dot-Product Attention
Vision-Language-Action (VLA)
Motion Structures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time-Frequency Geometric Cross-Attention
stationary wavelet transform
dot-product and wedge-product fusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shengye Dong
H
Haochen Niu
H
Hao Liu
P
Peiwen Lin
C
Chuang Wang
S
Shanmin Pang