When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistency between chain-of-thought (CoT) reasoning and execution, as well as the lack of safety monitoring, in vision-language-action (VLA) policies. It formally defines reasoning correctability and actionability for the first time, proposing the TRUST model. By integrating an offline value model, token-level reward mechanisms, and runtime steering techniques, TRUST monitors and corrects the reasoning generation of frozen VLAs in real time. Furthermore, this work pioneers the quantification of CoT's impact on embodied behavior, revealing that improved reasoning correctness does not necessarily yield task performance gains. Experiments demonstrate that TRUST reduces Alpamayo's collision rate by 30.4% and significantly enhances DeepThinkVLA's reasoning accuracy, validating the effectiveness of the proposed dual-axis evaluation framework.
📝 Abstract
Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning correctness from partial prefixes and uses these estimates to monitor and selectively steer reasoning generation in frozen VLA policies. On the Alpamayo 1.5 driving VLA, TRUST monitors correctness with 88.9% accuracy and improves reasoning correctness from 75.9% to 90.0%. On a baseline-defined challenging subset in AlpaSim, TRUST reduces collision rate by 30.4% and maximum trajectory error by 11.5% relative to the unsteered policy, outperforming a compute-matched Best-of-4 baseline. On the DeepThinkVLA manipulation VLA, TRUST improves the correctness of grasp-state claims from 69.3% to 90.2% and action-choice claims from 68.8% to 85.9%, yet closed-loop task performance on LIBERO-Plus remains largely unchanged. Empirical analysis reveals intent-consistent behavioral effects in Alpamayo 1.5 but limited effects in DeepThinkVLA, helping interpret these different task-level outcomes. Together, our results show that gains in reasoning correctness do not automatically imply gains in embodied performance, motivating evaluation of correctability and actionability when using CoT as a runtime safety interface.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action policies
Chain-of-Thought reasoning
runtime safety
correctability
actionability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought Steering
Vision-Language-Action Policies
Token-level Value Model
Runtime Safety Monitoring
Correctability and Actionability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sathwik Karnik
Safe and Intelligent Autonomy Lab, Stanford University
J
Joseph JR. Lee
Safe and Intelligent Autonomy Lab, Stanford University
Aryaman Gupta
Aryaman Gupta
PhD Student, Stanford University
RoboticsArtificial Intelligence
Somil Bansal
Somil Bansal
Assistant Professor, Stanford University
RoboticsArtificial intelligenceDynamic systems and control