Institution profile

Valeo

Industry researcheurope · fr
Official website
Research library25linked papers
Opportunities0open roles
Selected work

Representative Papers

CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving

Oct 08, 2026

Camera-only autonomous driving is inherently limited by occlusion and monocular depth uncertainty. This work proposes a Bayesian collaborative perception framework that leverages a VGGT feedforward network to generate uncertainty-aware 3D Gaussian representations, effectively modeling geometric ambiguity. To facilitate efficient multi-agent observation sharing over C-V2X communication, the method introduces dynamic object primitives (DOPs) requiring only 35 bytes each, thereby resolving depth ambiguity without relying on LiDAR. Extensive evaluations demonstrate that the proposed approach achieves performance improvements of 11.48% and 10.62% on the OPV2V+ and DAIR-V2X-C datasets, respectively, significantly outperforming existing vision-only methods.

0 citationsRead paper

Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

Oct 01, 2026

This work addresses the limited predictability of latent spaces in existing world models, which stems from the decoupling of representation learning and dynamics prediction. We propose an end-to-end joint training framework that integrates vision foundation models with flow matching generative models to synergistically optimize the latent encoder and the generative dynamics model, thereby directly shaping representations amenable to temporal prediction. Furthermore, a collapse-prevention mechanism is introduced to eliminate the reliance on two-stage training pipelines. The proposed approach significantly enhances long-horizon temporal coherence, consistently outperforming existing baselines across multi-task and high-resolution scenarios.

0 citationsRead paper

Pictura: Perspective-View Self-Play at Scale for Driving

Jul 28, 2026

This work addresses the representational gap that arises when driving policies trained with privileged state information are deployed using only first-person visual inputs. To bridge this gap, the authors propose the first purely vision-based self-play training framework that learns driving policies end-to-end directly from agent-centric images. Leveraging the GPU-accelerated multi-agent simulator Pictura and the PPO algorithm, the method achieves highly efficient training—processing 500,000 agent steps (equivalent to 2 million images) per second on a single H100 GPU. The resulting policy, Alberti, trained over 50 billion agent steps (approximately 35 million kilometers), closely matches the performance of privileged-observation baselines and demonstrates zero-shot superiority on re-rendered Waymo scenarios, effectively closing the perception gap between simulation and real-world deployment.

0 citationsRead paper

Human-like autonomy emerges from self-play and a pinch of human data

Jun 11, 2026

This work addresses the incompatibility between existing self-play reinforcement learning–derived driving policies and human driving behavior, which hinders effective coordination in real-world traffic despite ensuring safety. The authors propose a method that integrates an extremely small amount of human driving data—only 30 minutes, representing a 2,500-fold reduction compared to typical imitation learning—as a regularization target within a self-play reinforcement learning framework. This approach guides the policy to simultaneously achieve safety and mimic human driving styles, without requiring complex reward engineering or domain randomization. Trained on a single consumer-grade GPU in under 15 hours, the resulting policy demonstrates strong coordination capabilities with unseen human driving trajectories. Code and demonstration videos are publicly released.

0 citationsRead paper

Geometry-Aware Reinforcement Learning for 2D Irregular Nesting

Jun 09, 2026

This work addresses the limitations of traditional two-dimensional irregular nesting methods, which often suffer from insufficient geometric awareness and reliance on inefficient brute-force search strategies. To overcome these challenges, the authors propose a data-driven approach that integrates a geometry-aware neural encoder with reinforcement learning. The core innovation lies in the design of a Polygon Transformer (PoT) architecture featuring a cross-polygon attention mechanism, alongside the introduction of the first open-source training and evaluation benchmark tailored for complex geometric contours. Trained within a Combinatorial Optimization via Reinforcement Learning (CORL) framework, the resulting agent achieves area utilization performance on par with Sparrow, the current state-of-the-art heuristic solver, while significantly improving exploration efficiency in continuous placement spaces.

0 citationsRead paper
Recent publications

Latest Papers

CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving

Oct 08, 2026

Camera-only autonomous driving is inherently limited by occlusion and monocular depth uncertainty. This work proposes a Bayesian collaborative perception framework that leverages a VGGT feedforward network to generate uncertainty-aware 3D Gaussian representations, effectively modeling geometric ambiguity. To facilitate efficient multi-agent observation sharing over C-V2X communication, the method introduces dynamic object primitives (DOPs) requiring only 35 bytes each, thereby resolving depth ambiguity without relying on LiDAR. Extensive evaluations demonstrate that the proposed approach achieves performance improvements of 11.48% and 10.62% on the OPV2V+ and DAIR-V2X-C datasets, respectively, significantly outperforming existing vision-only methods.

0 citationsRead paper

Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

Oct 01, 2026

This work addresses the limited predictability of latent spaces in existing world models, which stems from the decoupling of representation learning and dynamics prediction. We propose an end-to-end joint training framework that integrates vision foundation models with flow matching generative models to synergistically optimize the latent encoder and the generative dynamics model, thereby directly shaping representations amenable to temporal prediction. Furthermore, a collapse-prevention mechanism is introduced to eliminate the reliance on two-stage training pipelines. The proposed approach significantly enhances long-horizon temporal coherence, consistently outperforming existing baselines across multi-task and high-resolution scenarios.

0 citationsRead paper

Pictura: Perspective-View Self-Play at Scale for Driving

Jul 28, 2026

This work addresses the representational gap that arises when driving policies trained with privileged state information are deployed using only first-person visual inputs. To bridge this gap, the authors propose the first purely vision-based self-play training framework that learns driving policies end-to-end directly from agent-centric images. Leveraging the GPU-accelerated multi-agent simulator Pictura and the PPO algorithm, the method achieves highly efficient training—processing 500,000 agent steps (equivalent to 2 million images) per second on a single H100 GPU. The resulting policy, Alberti, trained over 50 billion agent steps (approximately 35 million kilometers), closely matches the performance of privileged-observation baselines and demonstrates zero-shot superiority on re-rendered Waymo scenarios, effectively closing the perception gap between simulation and real-world deployment.

0 citationsRead paper

Human-like autonomy emerges from self-play and a pinch of human data

Jun 11, 2026

This work addresses the incompatibility between existing self-play reinforcement learning–derived driving policies and human driving behavior, which hinders effective coordination in real-world traffic despite ensuring safety. The authors propose a method that integrates an extremely small amount of human driving data—only 30 minutes, representing a 2,500-fold reduction compared to typical imitation learning—as a regularization target within a self-play reinforcement learning framework. This approach guides the policy to simultaneously achieve safety and mimic human driving styles, without requiring complex reward engineering or domain randomization. Trained on a single consumer-grade GPU in under 15 hours, the resulting policy demonstrates strong coordination capabilities with unseen human driving trajectories. Code and demonstration videos are publicly released.

0 citationsRead paper

Geometry-Aware Reinforcement Learning for 2D Irregular Nesting

Jun 09, 2026

This work addresses the limitations of traditional two-dimensional irregular nesting methods, which often suffer from insufficient geometric awareness and reliance on inefficient brute-force search strategies. To overcome these challenges, the authors propose a data-driven approach that integrates a geometry-aware neural encoder with reinforcement learning. The core innovation lies in the design of a Polygon Transformer (PoT) architecture featuring a cross-polygon attention mechanism, alongside the introduction of the first open-source training and evaluation benchmark tailored for complex geometric contours. Trained within a Combinatorial Optimization via Reinforcement Learning (CORL) framework, the resulting agent achieves area utilization performance on par with Sparrow, the current state-of-the-art heuristic solver, while significantly improving exploration efficiency in continuous placement spaces.

0 citationsRead paper