Institution profile

AI^2 Robotics

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

RFPO: Rectified Flow Policy Optimization for Embodied Control

Oct 07, 2026

This study addresses the degradation in control performance of flow policies under few-step execution caused by discretization errors, formally defining and resolving this "few-step discretization gap" for the first time. To this end, we propose the Rectified Flow Policy Optimization (RFPO) framework. RFPO corrects generation trajectories via reward-aware online reflow and integrates a Gaussian PPO controller for supervised optimization, enabling efficient and reliable robotic control under single-step Euler integration. Experimental results demonstrate that RFPO retains 98.5% of the full-step reward on the Go2 quadruped robot with single-step execution, reducing inference latency to 0.08 ms and achieving a 54.9× speedup. Furthermore, real-world deployment validates its stable control performance.

0 citationsRead paper

DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

Sep 30, 2026

This study addresses the inherent conflict between exploration noise and precise control in reinforcement learning for dexterous manipulation by proposing an explicit exploration scale scheduling strategy. By modeling exploration intensity as an annealing function of training steps, the method enables a smooth broad-to-narrow transition that decouples exploration from control, alongside a terminal task success-based evaluation criterion. Experiments integrating PPO, GRPO, and flow-parameterized FPO algorithms are conducted on the YCB dataset using a RealMan robotic arm and an Inspire dexterous hand. Results demonstrate that the proposed approach significantly improves simulation success rates. In real-world experiments, FPO increases grasping success from 25% to 85%, thoroughly validating the effectiveness of the exploration scheduling mechanism.

0 citationsRead paper

DeltaWAM: Delta World Action Models for Bimanual Manipulation

Sep 23, 2026

This study addresses the limitations of existing world-action models, specifically their prediction of redundant visual frames and high inference latency in bimanual robotic manipulation. To this end, it proposes a streaming incremental world-action model that jointly predicts visual deltas and actions rather than generating full frames, coupled with a streaming incremental memory mechanism to substantially reduce the computational overhead of video experts. Architecturally, the framework employs a synergistic multi-architecture design integrating dense anchors, sparse deltas, and action streams. Experimental results demonstrate that the proposed method improves task success rates to 85.4% while reducing training FLOPs by approximately 23% and inference latency by 36%, thereby achieving simultaneous optimization of both manipulation efficiency and computational performance.

0 citationsRead paper
Recent publications

Latest Papers

RFPO: Rectified Flow Policy Optimization for Embodied Control

Oct 07, 2026

This study addresses the degradation in control performance of flow policies under few-step execution caused by discretization errors, formally defining and resolving this "few-step discretization gap" for the first time. To this end, we propose the Rectified Flow Policy Optimization (RFPO) framework. RFPO corrects generation trajectories via reward-aware online reflow and integrates a Gaussian PPO controller for supervised optimization, enabling efficient and reliable robotic control under single-step Euler integration. Experimental results demonstrate that RFPO retains 98.5% of the full-step reward on the Go2 quadruped robot with single-step execution, reducing inference latency to 0.08 ms and achieving a 54.9× speedup. Furthermore, real-world deployment validates its stable control performance.

0 citationsRead paper

DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

Sep 30, 2026

This study addresses the inherent conflict between exploration noise and precise control in reinforcement learning for dexterous manipulation by proposing an explicit exploration scale scheduling strategy. By modeling exploration intensity as an annealing function of training steps, the method enables a smooth broad-to-narrow transition that decouples exploration from control, alongside a terminal task success-based evaluation criterion. Experiments integrating PPO, GRPO, and flow-parameterized FPO algorithms are conducted on the YCB dataset using a RealMan robotic arm and an Inspire dexterous hand. Results demonstrate that the proposed approach significantly improves simulation success rates. In real-world experiments, FPO increases grasping success from 25% to 85%, thoroughly validating the effectiveness of the exploration scheduling mechanism.

0 citationsRead paper

DeltaWAM: Delta World Action Models for Bimanual Manipulation

Sep 23, 2026

This study addresses the limitations of existing world-action models, specifically their prediction of redundant visual frames and high inference latency in bimanual robotic manipulation. To this end, it proposes a streaming incremental world-action model that jointly predicts visual deltas and actions rather than generating full frames, coupled with a streaming incremental memory mechanism to substantially reduce the computational overhead of video experts. Architecturally, the framework employs a synergistic multi-architecture design integrating dense anchors, sparse deltas, and action streams. Experimental results demonstrate that the proposed method improves task success rates to 85.4% while reducing training FLOPs by approximately 23% and inference latency by 36%, thereby achieving simultaneous optimization of both manipulation efficiency and computational performance.

0 citationsRead paper