🤖 AI Summary
This study addresses the failure of mainstream world model control methods, such as TD-MPC2 and DreamerV3, caused by action semantic mismatches arising from asynchronous execution. To resolve this, we elucidate the specific failure mechanisms inherent in these architectures and propose a lightweight execution-consistency interface. By leveraging temporal alignment and future sequence-action feedback to correct attribution errors, this approach adapts seamlessly to existing models without requiring retraining. Extensive evaluations across multi-domain simulations and real-world network tracing tasks validate both the accuracy of our diagnostic framework and the effectiveness of the proposed remediation strategy. Ultimately, this work provides a generalizable and computationally efficient solution for the asynchronous deployment of world models in practical control scenarios.
📝 Abstract
World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the action semantics assumed within world-model controllers, rather than treating it only as an external control disturbance. Through controlled interventions, we identify two architecture-dependent failure modes: planning-based controllers such as TD-MPC2 suffer from a future-action timeline mismatch between imagined and executed action sequences, while recurrent world models such as DreamerV3 can attribute observed transitions to commands that were not actually applied. Our analysis shows that TD-MPC2 requires the correct future action sequence during latent dynamics rollout, whereas DreamerV3 requires timely attribution of each transition to the action that generated it. Based on these findings, we introduce two lightweight execution-consistent interfaces, Future-Sequence for TD-MPC2 and Applied-Action Feedback for DreamerV3, that correct these mismatches without modifying the pretrained world models. Experiments across delays, packet loss, reordering, multiple control domains, measured network traces, and a process-separated asynchronous stack consistently support both diagnoses and the corresponding architecture-specific corrections.