ME-VLM:A Unified VLM for Embodied Cognition and Agent Coordination

๐Ÿ“… 2026-09-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บME-VLMๆจกๅž‹๏ผŒ้€š่ฟ‡็ป“ๅˆ็‰ฉ็†ๆ„Ÿ็Ÿฅใ€ๆ—ถ็ฉบๆŽจ็†ๅ’Œๅคšๆจกๆ€ไปฃ็†่ƒฝๅŠ›๏ผŒ่งฃๅ†ณไบ†็‰ฉ็†AIไธญ่ง†่ง‰่ฏญ่จ€็†่งฃไธŽ็Žฏๅขƒ็บฆๆŸๆ‰ง่กŒๅ้ฆˆ็š„่žๅˆ้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and execution feedback. We introduce MachEmbodied-VLM (ME-VLM), a unified vision-language model with two variants, 4B and 35B-A3B, that brings together embodied cognition and multimodal agent capabilities. Our work emphasizes physical perception and spatiotemporal reasoning, together with planning, interaction, and outcome assessment in both digital and physical environments. We construct training data spanning embodied and multimodal agent tasks, including execution observations and feedback to support outcome assessment and decision refinement. The training pipeline comprises embodied capability injection, separate reinforcement learning of embodied and multimodal-agent experts, and multi-teacher on-policy distillation that consolidates their complementary capabilities into a single model. Experiments show competitive performance on both embodied and agent benchmarks, as well as on autonomous-driving and embodied-navigation tasks. For edge deployment, visual token compression, W4A8 quantization, and hardware--software co-optimization enable on-device inference of the 4B variant on the M100, reducing prefill latency from 400 ms to 188 ms. Project Page: https://machembodied.com/ME-Brain/ME-VLM.html Code Repository: https://github.com/MachEmbodied/ME-VLM
Problem

Research questions and friction points this paper is trying to address.

Physical AI
embodied cognition
multimodal agent
environmental constraints
execution feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Vision-Language Model
Embodied Cognition
Multimodal Agent Coordination
On-Device Inference Optimization
๐Ÿ”Ž Similar Papers
2024-07-09IEEE/ASME transactions on mechatronicsCitations: 94
2024-10-04International Conference on Learning RepresentationsCitations: 0
F
Foundation Model
Li Auto Inc.
L
Li Auto Inc.
Li Auto Inc.