Institution profile

Wayve

Industry researcheurope · gb
Official website
Research library13linked papers
Opportunities96open roles
Selected work

Representative Papers

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration

May 21, 2026

This work addresses the challenge of sparse rewards and long-horizon exploration in 3D environments, where existing curiosity-driven methods often fall into local loops or redundant exploration. The authors propose a novel approach that integrates an online 3D reconstruction-based world model with RGB sequence-based policy learning, introducing spatial persistence and episodic context into the curiosity mechanism for the first time. This enables the agent to distinguish genuinely novel regions and avoid revisiting uninformative areas. Relying solely on visual inputs, the method achieves zero-shot generalization from Habitat-Matterport 3D (HM3D) to both Gibson and AI-generated environments after pure curiosity-driven pre-training, significantly outperforming from-scratch baselines and demonstrating strong performance on downstream tasks such as apple picking and image-goal navigation.

0 citationsRead paper

LA-Pose: Latent Action Pretraining Meets Pose Estimation

Apr 30, 2026

This work addresses the heavy reliance of camera pose estimation on large-scale 3D-annotated data by proposing a self-supervised pretraining approach that, for the first time, integrates inverse and forward dynamics models into this task. By learning latent action representations from large-scale unlabeled driving videos and using them as input features to a fine-tuned pose estimator, the method substantially reduces dependence on annotated data. With only a small amount of high-quality 3D annotations, it achieves over a 10% improvement in pose accuracy compared to state-of-the-art feedforward methods on both the Waymo and PandaSet benchmarks, attaining leading performance with an order of magnitude fewer annotations.

0 citationsRead paper

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

Apr 28, 2026

Existing 3D scene reconstruction methods struggle to meet the demands of interactive graphics applications for high-quality, cross-view consistent semantic segmentation. This work proposes a novel approach based on Radiant Foam—a voxelized Voronoi grid—by introducing an explicit semantic feature field at the cell level and, for the first time, directly incorporating spatial regularization into this field to achieve joint spatial-semantic decomposition. This design significantly enhances cross-view consistency and effectively mitigates artifacts caused by occlusions and inconsistent supervision. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches such as Gaussian Grouping and SAGA in both object-level semantic segmentation accuracy and consistency.

0 citationsRead paper
Recent publications

Latest Papers

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration

May 21, 2026

This work addresses the challenge of sparse rewards and long-horizon exploration in 3D environments, where existing curiosity-driven methods often fall into local loops or redundant exploration. The authors propose a novel approach that integrates an online 3D reconstruction-based world model with RGB sequence-based policy learning, introducing spatial persistence and episodic context into the curiosity mechanism for the first time. This enables the agent to distinguish genuinely novel regions and avoid revisiting uninformative areas. Relying solely on visual inputs, the method achieves zero-shot generalization from Habitat-Matterport 3D (HM3D) to both Gibson and AI-generated environments after pure curiosity-driven pre-training, significantly outperforming from-scratch baselines and demonstrating strong performance on downstream tasks such as apple picking and image-goal navigation.

0 citationsRead paper

LA-Pose: Latent Action Pretraining Meets Pose Estimation

Apr 30, 2026

This work addresses the heavy reliance of camera pose estimation on large-scale 3D-annotated data by proposing a self-supervised pretraining approach that, for the first time, integrates inverse and forward dynamics models into this task. By learning latent action representations from large-scale unlabeled driving videos and using them as input features to a fine-tuned pose estimator, the method substantially reduces dependence on annotated data. With only a small amount of high-quality 3D annotations, it achieves over a 10% improvement in pose accuracy compared to state-of-the-art feedforward methods on both the Waymo and PandaSet benchmarks, attaining leading performance with an order of magnitude fewer annotations.

0 citationsRead paper

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

Apr 28, 2026

Existing 3D scene reconstruction methods struggle to meet the demands of interactive graphics applications for high-quality, cross-view consistent semantic segmentation. This work proposes a novel approach based on Radiant Foam—a voxelized Voronoi grid—by introducing an explicit semantic feature field at the cell level and, for the first time, directly incorporating spatial regularization into this field to achieve joint spatial-semantic decomposition. This design significantly enhances cross-view consistency and effectively mitigates artifacts caused by occlusions and inconsistent supervision. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches such as Gaussian Grouping and SAGA in both object-level semantic segmentation accuracy and consistency.

0 citationsRead paper