🤖 AI Summary
This study addresses the absence of efficient, differentiable simulators for human articulated motion by proposing the first learnable articulated-body simulator based on a selective state space model (Mamba2). This work pioneers the integration of state space model architectures into physics simulation, predicting next-frame whole-body states using only positions, rotations, and action histories without requiring velocity inputs. Efficient inference is achieved through a novel rolling training protocol from scratch combined with CUDA graph compilation optimizations. Experimental results demonstrate that the proposed model attains 9334 FPS with sub-1% latency on an H100 GPU, enabling high-fidelity long-horizon trajectory prediction. It significantly outperforms conventional methods and supports seamless integration with video-based mesh recovery pipelines.
📝 Abstract
We introduce the first learned articulated body simulator based on a selective state space model (SSM), called PhysMamba. PhysMamba predicts next-frame full-body state from position, rotation, and joint-action history, without velocity inputs. We compare four architectures under partial- and full-observation inputs and three training protocols. The from-scratch rollout training protocol gives Mamba2 strong short- and mid-horizon accuracy under partial observation (s10 = 43 mm, 2/50 diverged), while the two-stage teacher-based rollout protocol stabilizes GRU but fails for Mamba2. With CUDA graph compilation, Mamba2 reaches 0.107 ms per frame (9,334 FPS) on an H100 GPU, within 1.1$\times$ of GRU's un-compiled throughput, adding under 1% latency to a 30 Hz HMR pipeline and enabling integration as a differentiable physics module for video-based mesh recovery.