Institution profile

Physical Intelligence

Industry researchnorthamerica · us
Official website
Research library9linked papers
Opportunities8open roles
Selected work

Representative Papers

EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning

Oct 06, 2026

This study addresses the challenge that first-person human data cannot be directly applied to robot control due to embodiment discrepancies, proposing the EgoLAP framework. Its core innovation lies in replacing low-level actions with cross-embodiment motion intentions and jointly learning human and robot trajectories through a shared language-action chain-of-thought. Furthermore, it introduces a motion-level reasoning mechanism integrating scene geometry, physics, and object affordances. Built upon vision-language-action (VLA) pretraining, this approach constructs a multimodal motion reasoning model. Real-world experiments demonstrate that EgoLAP achieves an average task progress of 80.1%, outperforming alternative action representations by a factor of 2.3 and surpassing composite reasoning formats.

0 citationsRead paper

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

Jun 09, 2026

Existing reinforcement learning methods often suffer from training instability and poor scalability in high-dimensional action spaces. This work proposes the QGF algorithm, which uniquely performs policy optimization entirely at test time: it first pretrains a highly expressive flow-based policy via behavioral cloning and then, during inference, leverages value gradients from a critic to guide the flow model toward generating higher-value actions—eliminating the need for additional policy training. Evaluated across multiple offline reinforcement learning benchmarks, QGF significantly outperforms existing test-time optimization approaches, matching or surpassing state-of-the-art training-time algorithms while incurring lower computational overhead, thereby achieving a favorable balance between stability and efficiency.

0 citationsRead paper

$π$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial Manipulation

Mar 26, 2026

This work addresses the challenge of transferring vision–language–action (VLA) models, pretrained on fixed-base robotic platforms, to highly dynamic and underactuated aerial manipulation systems, where dynamics mismatch severely degrades performance. To bridge this gap, the authors propose a payload-aware guidance mechanism that injects physical constraints during inference, alongside a synthetic navigation dataset generated via Gaussian splatting to alleviate real-world data scarcity. Notably, the approach enables effective zero-shot transfer to aerial grasping and navigation tasks without fine-tuning the base VLA model. Extensive real-world evaluation across 460 trials demonstrates that the synthetic data boosts navigation success from 81% to 100%, while payload-aware guidance increases grasping success from 23% to 50%. The integrated system achieves a 62% success rate on long-horizon compositional tasks.

0 citationsRead paper

MEM: Multi-Scale Embodied Memory for Vision Language Action Models

Mar 03, 2026

This work addresses the challenge of jointly modeling long-term semantic memory and short-term perceptual memory in end-to-end robot learning for complex, multi-stage tasks. The authors propose a multi-scale embodied memory architecture that, for the first time, integrates multimodal and multi-granularity memory mechanisms into robotic policy learning. Specifically, a video encoder compresses short-term visual memory to handle occlusions, while a language model processes long-term semantic memory represented in textual form. These components are unified within a vision–language–action policy framework. The approach successfully executes long-horizon tasks—such as kitchen cleaning and sandwich preparation—lasting up to fifteen minutes, demonstrating the ability to adapt manipulation strategies based on contextual cues.

0 citationsRead paper
Recent publications

Latest Papers

EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning

Oct 06, 2026

This study addresses the challenge that first-person human data cannot be directly applied to robot control due to embodiment discrepancies, proposing the EgoLAP framework. Its core innovation lies in replacing low-level actions with cross-embodiment motion intentions and jointly learning human and robot trajectories through a shared language-action chain-of-thought. Furthermore, it introduces a motion-level reasoning mechanism integrating scene geometry, physics, and object affordances. Built upon vision-language-action (VLA) pretraining, this approach constructs a multimodal motion reasoning model. Real-world experiments demonstrate that EgoLAP achieves an average task progress of 80.1%, outperforming alternative action representations by a factor of 2.3 and surpassing composite reasoning formats.

0 citationsRead paper

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

Jun 09, 2026

Existing reinforcement learning methods often suffer from training instability and poor scalability in high-dimensional action spaces. This work proposes the QGF algorithm, which uniquely performs policy optimization entirely at test time: it first pretrains a highly expressive flow-based policy via behavioral cloning and then, during inference, leverages value gradients from a critic to guide the flow model toward generating higher-value actions—eliminating the need for additional policy training. Evaluated across multiple offline reinforcement learning benchmarks, QGF significantly outperforms existing test-time optimization approaches, matching or surpassing state-of-the-art training-time algorithms while incurring lower computational overhead, thereby achieving a favorable balance between stability and efficiency.

0 citationsRead paper

$π$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial Manipulation

Mar 26, 2026

This work addresses the challenge of transferring vision–language–action (VLA) models, pretrained on fixed-base robotic platforms, to highly dynamic and underactuated aerial manipulation systems, where dynamics mismatch severely degrades performance. To bridge this gap, the authors propose a payload-aware guidance mechanism that injects physical constraints during inference, alongside a synthetic navigation dataset generated via Gaussian splatting to alleviate real-world data scarcity. Notably, the approach enables effective zero-shot transfer to aerial grasping and navigation tasks without fine-tuning the base VLA model. Extensive real-world evaluation across 460 trials demonstrates that the synthetic data boosts navigation success from 81% to 100%, while payload-aware guidance increases grasping success from 23% to 50%. The integrated system achieves a 62% success rate on long-horizon compositional tasks.

0 citationsRead paper

MEM: Multi-Scale Embodied Memory for Vision Language Action Models

Mar 03, 2026

This work addresses the challenge of jointly modeling long-term semantic memory and short-term perceptual memory in end-to-end robot learning for complex, multi-stage tasks. The authors propose a multi-scale embodied memory architecture that, for the first time, integrates multimodal and multi-granularity memory mechanisms into robotic policy learning. Specifically, a video encoder compresses short-term visual memory to handle occlusions, while a language model processes long-term semantic memory represented in textual form. These components are unified within a vision–language–action policy framework. The approach successfully executes long-horizon tasks—such as kitchen cleaning and sandwich preparation—lasting up to fifteen minutes, demonstrating the ability to adapt manipulation strategies based on contextual cues.

0 citationsRead paper