🤖 AI Summary
This study addresses the prohibitive per-step trajectory sampling costs of data-driven Model Predictive Control (MPC), which hinder real-time robotic control. Inspired by human dual-process cognition, this work proposes Fast-TD-MPC, a framework that introduces an adaptive deliberation mechanism to depart from the conventional paradigm of exhaustive planning at every step. By integrating a learned world model, online trajectory optimization, and a lightweight routing algorithm, the method enables on-demand switching between rapid policy execution and deep planning. Computationally intensive reasoning is invoked only at critical states, effectively balancing efficiency and performance. Evaluated across 103 continuous control tasks, Fast-TD-MPC achieves approximately fourfold inference acceleration while maintaining competitive performance, and preserves robustness under external perturbations through fallback planning.
📝 Abstract
Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts deployment to control frequencies well below what real-time robotics demands. Motivated by the dual-process theory of human cognition, which distinguishes between fast, intuitive processing (System 1) and slower, deliberative reasoning (System 2), we ask whether every decision requires the same degree of computational deliberation. We propose Fast-TD-MPC, a lightweight framework that adaptively routes between fast policy execution and test-time planning, reserving costly deliberation for states where it is most needed. Fast-TD-MPC delivers competitive task performance across 103 continuous control tasks while achieving up to ~4x faster inference. Under external disturbances, Fast-TD-MPC selectively falls back to planning, maintaining robustness comparable to the original planner.