General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

📅 2026-05-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a fundamental limitation in existing action modeling approaches, which regress absolute coordinates and thereby violate the principle of general covariance, leading to policies that are sensitive to motion style and speed and exhibit limited generalization. To overcome this, the paper introduces the Generalized Action Manifold (GAM) framework—the first to incorporate general covariance into embodied intelligence. GAM employs arc-length parameterization to decouple spatial trajectories from temporal dynamics and integrates mode-affine decomposition with a pose-normalized coordinate system, constructing a spatiotemporally orthogonal-invariant continuous action manifold. Within a structured vision–language–action architecture, this approach achieves complete disentanglement of geometric structure and temporal dynamics, enabling dense, effective policies to be generalized from sparse demonstrations. The method significantly outperforms baselines in cross-scenario and variable-speed tasks, demonstrating exceptional transferability and robustness.
📝 Abstract
Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordinates, which violates the principle of general covariance. Fundamentally, this conflates the intrinsic task geometry with rigid execution patterns, binding policies to specific motion styles and fixed speeds. To resolve this, we propose the Generalized Action Manifold (GAM) framework that enforces general covariance through structural disentanglement. Specifically, GAM realizes the manifold by enforcing invariance across two orthogonal dimensions: (1) Temporal Invariance, utilizing an Arc-Length Parameterizer to orthogonalize the spatial path geometry from temporal dynamics, ensuring robustness to velocity variations; (2) Geometric Invariance, where a Schema-Affine-Factorization mechanism maps trajectories to canonical ``world lines'' in a pose-normalized coordinate frame. This distinguishes invariant geometric schemas from affine modulations, ensuring spatial generalizability. By integrating GAM within a structured Vision-Language-Action (VLA) architecture, we enable sparse demonstrations to densely populate a continuous, valid action manifold. Empirical results demonstrate that GAM enables superior transfer and robustness capabilities, outperforming geometry-agnostic baselines.
Problem

Research questions and friction points this paper is trying to address.

general covariance
embodied intelligence
action generalization
spatio-temporal decoupling
geometric invariance
Innovation

Methods, ideas, or system contributions that make the work stand out.

General Covariance
Spatio-Temporal Decoupling
Action Manifold
Arc-Length Parameterization
Schema-Affine Factorization
🔎 Similar Papers
No similar papers found.