Where Do Embodied Decisions Come From? Rethinking Latent and Explicit Reasoning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the potential performance degradation caused by explicit chain-of-thought reasoning in embodied agents by proposing a "decide-then-explain" paradigm that decouples latent decision-making computations from explicit reasoning processes to optimize action generation. By introducing contribution metrics for visual and reasoning conditioning, this work reveals capability-dependency effects and redefines the role of reasoning during both training and inference. Leveraging vision-language-action models alongside intervention-based experimental analyses, the proposed method consistently outperforms conventional baselines across autonomous driving and robotic manipulation tasks. These results demonstrate that decoupling implicit policy computation from post-hoc verbalized reasoning significantly enhances perception-driven decision-making capabilities in embodied AI systems.
📝 Abstract
Chain-of-Thought (CoT) reasoning is increasingly incorporated into Vision-Language-Action (VLA) models, yet it can degrade the performance of stronger embodied agents. We investigate this capability-dependent effect by distinguishing explicit reasoning from latent decision computation, i.e., perception-grounded computation that directly supports action prediction. Under the standard \textit{think-then-act} (TTA) paradigm, intervening on the generated CoT while fixing the visual input and model parameters causes a substantial performance collapse, with Navigation F1 dropping from 72.14% to 11.84%, demonstrating the strong influence of explicit reasoning on action generation. We then propose \textit{decide-then-explain} (DTE), which predicts actions before generating explanations, and introduce Visual Conditional Contribution (VCC) and Reasoning Conditional Contribution (RCC) to characterize the resulting decision process. Across autonomous driving and robotic manipulation, DTE consistently outperforms TTA and conventional \textit{no-CoT} baselines, while exhibiting greater reliance on perception-grounded computation. Further TTA-trained, DTE-inference experiments show that this benefit is not solely attributable to retraining under the new factorization. Our results suggest that for capable embodied agents, explicit CoT may be better used to shape decision computation during training rather than mediate action generation at inference time. Code: https://github.com/ocean-luna/openvla-decide-then-explain.
Problem

Research questions and friction points this paper is trying to address.

Embodied AI
Chain-of-Thought
Vision-Language-Action models
Latent reasoning
Explicit reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decide-then-Explain (DTE)
Chain-of-Thought (CoT)
Vision-Language-Action (VLA)
Visual Conditional Contribution (VCC)
Embodied AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yuan Lin
Yuan Lin
Ocean College, Zhejiang University
RheologyPolymer physcisMulti-phase flow
Z
Ziyue Zhou
Li Auto Inc.
J
JinLong Zhao
Li Auto Inc.
P
Pei Liu
Li Auto Inc.
H
Haipeng Liu
Li Auto Inc.
P
Pan Zhou
Li Auto Inc.
K
Kun Zhan
Li Auto Inc.