🤖 AI Summary
This study addresses the limitation of conventional aggregate accuracy in inferring player inputs from gameplay videos, which often obscures recognition failures for rare actions. To overcome this, we propose a macro-F1-based, key-level fine-grained evaluation framework. In data-constrained scenarios, we construct large-scale inverse dynamics models that incorporate optical flow preprocessing to extract spatial motion features, validating the impact of architectures and training objectives across Trackmania and Cyberpunk 2077. Our findings reveal the critical roles of model architecture and motion representation while exposing inherent ambiguities arising from camera motion. Ultimately, this work highlights these challenges and points toward integrating explicit 3D structural modeling as a promising direction for optimizing future inverse dynamics performance.
📝 Abstract
Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as $F_1^{macro}$. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.