Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of existing dual-system Vision-Language-Action (VLA) models, which invoke slow reasoning modules at fixed frequencies, causing computational waste on simple tasks and response latency in complex scenarios. We propose TUD, an adaptive reasoning framework that quantifies uncertainty via the cross-step dispersion of action re-predictions, dynamically triggering a general-purpose reasoner without requiring additional annotations or auxiliary models, thereby achieving zero-overhead adaptive scheduling. Experimental evaluations on VLA-Arena and real-world robotic deployments demonstrate that our method reduces general-purpose reasoning invocations by 75% compared to the strongest baseline, while simultaneously improving task success rates and computational cost-effectiveness.
📝 Abstract
Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay renewed reasoning when the scene changes unexpectedly. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive and is computed from forwards the architecture already runs, requiring neither manual phase labels nor an auxiliary uncertainty model. On VLA-Arena, it achieves a higher success rate at matched call budgets than alternative uncertainty baselines while maintaining low wall-clock overhead, and more consistently separates successful from failed rollouts. Also, TUD finds a more favorable cost-success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. The same trade-off appears in our real-robot experiments, where TUD cuts generalist calls by 75% relative to the strongest fixed-interval baseline while achieving an even higher success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA)
Dual-system model
Adaptive inference
Predictive uncertainty
Robotic control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-System VLA
Predictive Uncertainty
Adaptive Inference
Action Re-prediction Dispersion
Efficient Robotic Control