Institution profile

Applied Intuition

Industry researchnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

STRIKE: Learning Visual State Transitions for Physical World Modeling

Oct 07, 2026

This study addresses the limitation of existing physical world modeling approaches, which predominantly focus on generating coherent motion while struggling to accurately capture interaction-induced scene state changes. To overcome this, we propose STRIKE, a framework that innovatively decouples visual state transition learning from dense video generation. Specifically, STRIKE employs an image-based transition model to predict subsequent scene configurations and integrates a pretrained vision-language planner to recursively generate future state sequences. A dynamics model then renders complete videos, with event-aligned supervision introduced to optimize training. By leveraging visual state transitions as intermediate representations, the proposed method effectively disentangles state prediction from video rendering. Extensive evaluations on benchmarks such as Physics-IQ Verified demonstrate that STRIKE significantly outperforms existing video backbone baselines in both physical consistency and manipulation fidelity.

0 citationsRead paper

S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens

Oct 06, 2026

This study addresses the challenge of maintaining persistent scene states in streaming 3D reconstruction to enable online fusion and real-time rendering. To this end, it proposes a feedforward framework based on adaptive-sized persistent tokens. Specifically, a spatial-aware Transformer is employed to integrate new observations with historical states, while a selective admission module distinguishes between scene updates and expansions. Furthermore, a hierarchical decoder directly generates 3D Gaussians without requiring the caching of historical frames. Evaluated across four benchmarks, the proposed method achieves competitive streaming rendering quality through a compact Gaussian representation, effectively resolving the memory and efficiency bottlenecks inherent in online 3D reconstruction.

0 citationsRead paper

Building Rome from a Single Image

Oct 06, 2026

This study addresses the challenge of generating complete 3D scenes from a single image, where existing methods struggle to accommodate both indoor and outdoor environments while reconstructing geometry in unobserved regions. Building upon the Trellis 2 architecture, this work reformulates an object-centric generator by introducing a distance-aware adaptive chunking strategy to enhance the perception of free space and unobserved areas. Furthermore, explicit 2D-3D correspondence and feature lifting techniques are incorporated to ensure geometric consistency. To facilitate training, a large-scale synthetic outdoor dataset is constructed. Experimental results demonstrate that the proposed method significantly outperforms existing baselines across multiple benchmarks in terms of both geometric accuracy and perceptual quality, achieving high-fidelity 3D mesh generation for diverse indoor and outdoor scenes.

0 citationsRead paper

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

Sep 28, 2026

This study addresses the exploration bottleneck in reinforcement learning for manipulating multi-degree-of-freedom objects by proposing a policy training framework guided by demonstration-informed reset distributions. Rather than directly imitating actions, the method leverages human hand-object interaction demonstrations to construct reset distributions through kinematic retargeting and state filtering, thereby driving efficient exploration under a task-agnostic generic reward. The proposed approach successfully trains universal manipulation policies across three heterogeneous embodiments, achieving zero-shot cross-embodiment transfer and zero-shot sim-to-real deployment. By decoupling exploration guidance from action-level imitation, this work provides a highly generalizable solution for complex dexterous manipulation tasks.

0 citationsRead paper

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

May 22, 2026

Current video generation models often fail to maintain physical and motion consistency, limiting their reliability as world simulators. This work proposes a self-supervised latent motion prior that leverages only unlabeled videos to model inter-frame motion dynamics within the latent space of diffusion models, without requiring external supervision. By incorporating two lightweight components—a macroscopic motion drift loss and a microscopic motion field guidance—the method effectively enhances the physical plausibility of generated videos. Evaluated on VideoPhy and VideoPhy2 benchmarks, the approach outperforms baselines that rely on external supervisory signals, while maintaining competitive overall generation quality on VBench and achieving significant improvements in motion-related metrics.

0 citationsRead paper
Recent publications

Latest Papers

STRIKE: Learning Visual State Transitions for Physical World Modeling

Oct 07, 2026

This study addresses the limitation of existing physical world modeling approaches, which predominantly focus on generating coherent motion while struggling to accurately capture interaction-induced scene state changes. To overcome this, we propose STRIKE, a framework that innovatively decouples visual state transition learning from dense video generation. Specifically, STRIKE employs an image-based transition model to predict subsequent scene configurations and integrates a pretrained vision-language planner to recursively generate future state sequences. A dynamics model then renders complete videos, with event-aligned supervision introduced to optimize training. By leveraging visual state transitions as intermediate representations, the proposed method effectively disentangles state prediction from video rendering. Extensive evaluations on benchmarks such as Physics-IQ Verified demonstrate that STRIKE significantly outperforms existing video backbone baselines in both physical consistency and manipulation fidelity.

0 citationsRead paper

S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens

Oct 06, 2026

This study addresses the challenge of maintaining persistent scene states in streaming 3D reconstruction to enable online fusion and real-time rendering. To this end, it proposes a feedforward framework based on adaptive-sized persistent tokens. Specifically, a spatial-aware Transformer is employed to integrate new observations with historical states, while a selective admission module distinguishes between scene updates and expansions. Furthermore, a hierarchical decoder directly generates 3D Gaussians without requiring the caching of historical frames. Evaluated across four benchmarks, the proposed method achieves competitive streaming rendering quality through a compact Gaussian representation, effectively resolving the memory and efficiency bottlenecks inherent in online 3D reconstruction.

0 citationsRead paper

Building Rome from a Single Image

Oct 06, 2026

This study addresses the challenge of generating complete 3D scenes from a single image, where existing methods struggle to accommodate both indoor and outdoor environments while reconstructing geometry in unobserved regions. Building upon the Trellis 2 architecture, this work reformulates an object-centric generator by introducing a distance-aware adaptive chunking strategy to enhance the perception of free space and unobserved areas. Furthermore, explicit 2D-3D correspondence and feature lifting techniques are incorporated to ensure geometric consistency. To facilitate training, a large-scale synthetic outdoor dataset is constructed. Experimental results demonstrate that the proposed method significantly outperforms existing baselines across multiple benchmarks in terms of both geometric accuracy and perceptual quality, achieving high-fidelity 3D mesh generation for diverse indoor and outdoor scenes.

0 citationsRead paper

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

Sep 28, 2026

This study addresses the exploration bottleneck in reinforcement learning for manipulating multi-degree-of-freedom objects by proposing a policy training framework guided by demonstration-informed reset distributions. Rather than directly imitating actions, the method leverages human hand-object interaction demonstrations to construct reset distributions through kinematic retargeting and state filtering, thereby driving efficient exploration under a task-agnostic generic reward. The proposed approach successfully trains universal manipulation policies across three heterogeneous embodiments, achieving zero-shot cross-embodiment transfer and zero-shot sim-to-real deployment. By decoupling exploration guidance from action-level imitation, this work provides a highly generalizable solution for complex dexterous manipulation tasks.

0 citationsRead paper

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

May 22, 2026

Current video generation models often fail to maintain physical and motion consistency, limiting their reliability as world simulators. This work proposes a self-supervised latent motion prior that leverages only unlabeled videos to model inter-frame motion dynamics within the latent space of diffusion models, without requiring external supervision. By incorporating two lightweight components—a macroscopic motion drift loss and a microscopic motion field guidance—the method effectively enhances the physical plausibility of generated videos. Evaluated on VideoPhy and VideoPhy2 benchmarks, the approach outperforms baselines that rely on external supervisory signals, while maintaining competitive overall generation quality on VBench and achieving significant improvements in motion-related metrics.

0 citationsRead paper