🤖 AI Summary
This work addresses the challenge of efficiently determining, prior to action execution, whether it is worthwhile to invest additional computation in test-time scaling for robotic control. The authors propose MethodGated, a lightweight, training-free framework that introduces zero-shot geometric consistency evaluation into test-time scaling of World Action Models (WAMs). By leveraging a frozen geometric foundation model, MethodGated scores the cross-view depth reprojection consistency between WAM-generated actions and their predicted future observations. Integrating a Best-of-N strategy with an action-future consistency gating mechanism, the framework dynamically decides whether to activate test-time scaling. Evaluated on benchmarks such as RoboCasa, MethodGated achieves 74.8% of the full performance gain while triggering scaling in only 26.2% of cases, substantially improving task success rates under fixed computational budgets.
📝 Abstract
Test-time scaling improves foundation-model inference by spending additional computation, but robot control requires deciding whether extra compute is useful before executing an action. World Action Models (WAMs) make this decision natural: each rollout exposes both an action chunk and predicted future observations. We propose \methodgated, a training-free selective test-time scaling framework for WAMs. We first instantiate \method, a fixed-budget Best-of-$N$ selector that ranks sampled rollouts by cross-view depth reprojection consistency of their predicted futures, computed with a frozen geometry foundation model. \methodgated\ adds a lightweight action--future consistency gate that invokes \method\ only when the initial rollout appears internally inconsistent. Across five benchmark--backbone settings on RoboCasa, LIBERO Long, and RoboTwin~2.0, fixed-budget \method\ improves $N{=}8$ task success in every setting, e.g., raising the RoboCasa group average from $66.3\%$ to $68.4\%$ with Cosmos Policy and from $80.8\%$ to $82.5\%$ with X-WAM. With gating enabled, \methodgated\ recovers on average $74.8\%$ of the always-on success gain while triggering additional sampling on only $26.2\%$ of decision points. Offline diagnostics show that cross-view reprojection is a strong task-label-free selector, and we identify false low-score selections as a failure mode that helps explain why performance can saturate or degrade as $N$ increases.