🤖 AI Summary
This study addresses the high inference costs of world action models, where the engineering complexity of combining existing acceleration techniques hinders scalable deployment. We propose an automated inference acceleration framework based on coding agents that employs a bottleneck-driven iterative optimization workflow. By synergizing performance profiling, verification tools, and hardware-aware optimization techniques, the framework enables automatic search and deployment of acceleration strategies. Experimental results demonstrate that our approach achieves up to 9.95× inference speedup while preserving action generation quality and task success rates without degradation, significantly reducing system latency. This work establishes a general, automated paradigm for the efficient deployment of world action models.
📝 Abstract
World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95x lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.