EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the over-concentration of action computation within single experts and the lack of semantic-visual-motor synergy in unified embodied models by proposing the EWAM framework. This method introduces an asymmetric joint attention mechanism that integrates cross-embodiment trajectories with first-person videos during pretraining, validated through causal intervention, to enable a sequential progression from semantic understanding and visual foresight to action generation. Without hierarchical supervision, the model spontaneously develops depth-wise specialization: shallow layers focus on vision-language features, intermediate layers predict future frames, and deep layers specialize in action generation. Simulation and real-world robotic experiments demonstrate that EWAM significantly outperforms existing baselines, effectively enhancing cross-embodiment transferability, robustness, and long-horizon task completion rates.
📝 Abstract
Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embodied model whose asymmetric joint attention lets action tokens read semantic, current-visual, predicted-future, and action information at every layer while the perceptual experts retain their distinct roles. Without layer-wise supervision, EWAM develops an emergent depth-wise specialization: action queries attend mainly to vision-language features in shallow layers, to predicted future frames in intermediate layers, and to action tokens themselves in deep layers. This handoff replicates across tasks and is stable across denoising steps. Checkpoint tracking and causal interventions show that it is learned and that action generation depends on it. EWAM is pretrained in two separate regimes, one on cross-embodiment robot trajectories and one on human egocentric video. In simulation and real-robot experiments, it surpasses existing VLA, WAM, and hybrid baselines. Human egocentric data improve both cross-embodiment transfer and real-robot robustness, and subtask-phase supervision improves long-horizon completion. Together, these results suggest that unified embodied learning can induce an ordered internal progression from semantic understanding, through visual foresight, to action formation.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
World-Action Models
Embodied AI
Unified Model
Action Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
World-Action Model
Asymmetric Joint Attention
Emergent Specialization
Unified Embodied Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.