Institution profile

Beijing XYZ embodied AI Co., Ltd.

Industry researchasia · cn
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Sep 29, 2026

This study addresses the reasoning failures of vision-language models in long-horizon robotic manipulation caused by the absence of local physical interaction feedback. To overcome this limitation, we propose a dual-loop framework that innovatively introduces Hierarchical Physical Knowledge (HPK), transforming repetitive physical interactions into retrievable and correctable structured knowledge assets to enable self-improvement without updating the base model. By integrating knowledge representation, physical feedback revision, and scene grounding techniques, our approach achieves an average success rate improvement of 24.2% on RMBench. Notably, on the GPT-5.5 hold-out set, the success rate increases substantially from 48.3% to 75.0%, while demonstrating significant zero-shot cross-benchmark transferability.

0 citationsRead paper

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Aug 10, 2026

This work addresses the high deployment cost of existing video-based generative world models and the suboptimal control performance caused by decoupling state prediction from action modeling in latent approaches. It introduces, for the first time, the Joint-Embedding Predictive Architecture (V-JEPA) to world modeling, constructing a shared predictor in the pretrained V-JEPA latent space that jointly learns state transitions and continuous action generation. A structured current–future joint objective preserves dense visual correspondences, enabling tight coupling between prediction and policy. The resulting model integrates seamlessly into vision–language–action (VLA) systems. Evaluated on LIBERO-Plus, it achieves a 79.2% success rate—the best among methods without large-scale robot pretraining—and reaches a new state-of-the-art overall performance of 86.3% with the π₀.₅ instantiation, while demonstrating strong generalization on RoboTwin 2.0 and real-world dual-arm tasks.

0 citationsRead paper
Recent publications

Latest Papers

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Sep 29, 2026

This study addresses the reasoning failures of vision-language models in long-horizon robotic manipulation caused by the absence of local physical interaction feedback. To overcome this limitation, we propose a dual-loop framework that innovatively introduces Hierarchical Physical Knowledge (HPK), transforming repetitive physical interactions into retrievable and correctable structured knowledge assets to enable self-improvement without updating the base model. By integrating knowledge representation, physical feedback revision, and scene grounding techniques, our approach achieves an average success rate improvement of 24.2% on RMBench. Notably, on the GPT-5.5 hold-out set, the success rate increases substantially from 48.3% to 75.0%, while demonstrating significant zero-shot cross-benchmark transferability.

0 citationsRead paper

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Aug 10, 2026

This work addresses the high deployment cost of existing video-based generative world models and the suboptimal control performance caused by decoupling state prediction from action modeling in latent approaches. It introduces, for the first time, the Joint-Embedding Predictive Architecture (V-JEPA) to world modeling, constructing a shared predictor in the pretrained V-JEPA latent space that jointly learns state transitions and continuous action generation. A structured current–future joint objective preserves dense visual correspondences, enabling tight coupling between prediction and policy. The resulting model integrates seamlessly into vision–language–action (VLA) systems. Evaluated on LIBERO-Plus, it achieves a 79.2% success rate—the best among methods without large-scale robot pretraining—and reaches a new state-of-the-art overall performance of 86.3% with the π₀.₅ instantiation, while demonstrating strong generalization on RoboTwin 2.0 and real-world dual-arm tasks.

0 citationsRead paper