🤖 AI Summary
This study addresses the limited task adaptability of robotic policies caused by fixed execution step sizes by modeling the step size as a latent variable and proposing a training-free adaptive framework based on dynamic inference from action expert evidence. The core methodology introduces a training-free AHS selector and an optional QHA adapter, which leverage spectral stability and continuity evidence to achieve adaptive planning. To further enhance inference robustness, the framework incorporates Beta posterior tracking, a kernel forgetting mechanism, and query-based attention fusion. Experimental evaluations on the RoboTwin and RoboCasa benchmarks demonstrate that the proposed approach significantly improves task success rates, achieving gains of up to 9.67%. Moreover, real-world deployments confirm its superior performance in practical scenarios.
📝 Abstract
Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce ChunkTrust, which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware Horizon Selector (AHS) combines intra-chunk spectral stability of generation traces with inter-chunk continuity between executed history and predicted actions. An online Beta posterior with kernel forgetting tracks horizon preferences across replans. A lightweight Query-based Horizon Adapter (QHA) optionally learns a context-conditioned dense prior from complementary evidence, fused with current evidence and episode-local Beta memory while the base policy remains frozen. Across RoboTwin2.0 and RoboCasa GR1 Tabletop, AHS improves overall task-averaged success for each evaluated base-policy configuration, including gains of +6.80 percentage points on $π_{0.5}$ over all 50 RoboTwin2.0 tasks and +9.67 percentage points on Qwen3GR00T in RoboCasa. AHS+QHA raises the gain over Base to +9.44 percentage points on the eight-task $π_{0.5}$ evaluation. On four real-world household tasks, AHS improves the equal-task mean normalized process score from 50.4% to 57.5%. Ablations examine the contributions of both evidence terms, temporal memory, and the learned prior. Project page is https://hf618.github.io/ChunkTrust.github.io/