DualWAM: Dual-System World Action Models for Asynchronous Global Planning and Local Refinement
本文提出DualWAM方法,通过解耦全局规划和局部细化,解决未来视觉预测计算成本高和闭环响应性差的问题。
本文提出DualWAM方法,通过解耦全局规划和局部细化,解决未来视觉预测计算成本高和闭环响应性差的问题。
为解决机器人T恤折叠和展开难题,本文构建了一个大规模合成数据集FoldNet++,并通过基于规则的框架生成操作演示来训练视觉运动策略。
研究通过UniPart模型解决了零样本语言驱动的3D部分分割问题,利用CLIP文本嵌入和大规模标注数据集LangPart-1M进行训练。
本文提出RoboGesture框架,通过设计数据、建模和控制来解决人形机器人同步生成语义手势的问题,采用层次语义-声学对齐器和流式条件运动生成器等方法。
Existing evaluations of humanoid motion tracking rely primarily on kinematic errors, which fail to capture physically implausible distortions perceptible to humans—such as foot sliding or incorrect contact—and suffer from small-scale, low-diversity test sets. To address these limitations, this work introduces HumanTracker, a large-scale and diverse benchmark comprising 153 hours of optical motion capture data from professional actors, covering four action categories with accompanying textual annotations. Furthermore, the authors propose HumanScore, a human-aligned evaluation metric derived from a preference model trained on 12K motion pairs (24K individual motions). HumanScore enables fine-grained diagnosis of critical physical properties like contact fidelity and support stability, significantly outperforming conventional metrics across multiple state-of-the-art trackers, accurately predicting human preferences, and uncovering previously overlooked physical inconsistencies.
本文提出DualWAM方法,通过解耦全局规划和局部细化,解决未来视觉预测计算成本高和闭环响应性差的问题。
为解决机器人T恤折叠和展开难题,本文构建了一个大规模合成数据集FoldNet++,并通过基于规则的框架生成操作演示来训练视觉运动策略。
研究通过UniPart模型解决了零样本语言驱动的3D部分分割问题,利用CLIP文本嵌入和大规模标注数据集LangPart-1M进行训练。
本文提出RoboGesture框架,通过设计数据、建模和控制来解决人形机器人同步生成语义手势的问题,采用层次语义-声学对齐器和流式条件运动生成器等方法。
Existing evaluations of humanoid motion tracking rely primarily on kinematic errors, which fail to capture physically implausible distortions perceptible to humans—such as foot sliding or incorrect contact—and suffer from small-scale, low-diversity test sets. To address these limitations, this work introduces HumanTracker, a large-scale and diverse benchmark comprising 153 hours of optical motion capture data from professional actors, covering four action categories with accompanying textual annotations. Furthermore, the authors propose HumanScore, a human-aligned evaluation metric derived from a preference model trained on 12K motion pairs (24K individual motions). HumanScore enables fine-grained diagnosis of critical physical properties like contact fidelity and support stability, significantly outperforming conventional metrics across multiple state-of-the-art trackers, accurately predicting human preferences, and uncovering previously overlooked physical inconsistencies.