ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出ME-Dex-1.0模型,通过结合视觉、触觉和动作学习来预测未来状态,解决了机器人操作中多源异构触觉信号的融合问题。
📝 Abstract
World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.
Problem

Research questions and friction points this paper is trying to address.

Tactile Sensing
World Action Models
Predictive Capabilities
Future Observations
Joint Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Transformers
Shared Attention
Canonical Hand Model
Unified Tactile Autoencoder
Agentic Tactile Data Engine
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xuancheng Zhang
Xuancheng Zhang
Tsinghua University
3D VisionDeep Learning
X
Xuetao Liu
Q
Qianying Tang
J
Jizhe Wang
Z
Zhijing Cheng
B
Bochen Lin
H
Haoran Wen
M
Ming Li
K
Kun Zhan
Y
Yu Liu