Senior Context Fusion AI Engineer - Autonomous Vehicles

Nvidia
US, CA, Santa Clara / US, WA, Redmond / US, NV, Remote2026-09-18remote_local

About the job

We are looking for a strong engineer to join the DRIVE Road Structure / Online Mapping / Context Fusion team. In this role, you will help craft and guide the future of our L3/L4 autonomous-driving solution by building a complete, learned 3D/4D world model that fuses navigation, ego-motion, perception, and sensor signals. You will work closely with perception, prediction, planning, and simulation teams to deliver a world representation that is complete, temporally consistent, uncertainty-aware, and robust enough to drive through the most challenging roads and intersections in L3/L4 autonomy level.

Responsibilities

Design and develop learning-based, multimodal sensor-fusion systems that transform synchronized sensor history, ego-motion, navigation context, and driving context into a unified spatiotemporal world representation.\nBuild architectures that jointly reason over camera, LiDAR, radar, and vehicle-state inputs, with appropriate handling of calibration, synchronization, coordinate transforms, sensor latency, and uncertainty.\nDevelop end-to-end and multi-task models that produce driving-relevant outputs from a shared scene representation, including; road graph elements such as lanes, boundaries, crosswalks, and traffic controls; semantic scene understanding; occupancy and free-space representations, including uncertain and occluded regions;\nDevelop scalable multimodal fusion architectures, including Transformer-based early, late, and hierarchical fusion; BEV, point/voxel, and image-based representations; temporal context aggregation; and cross-modal attention.\nCreate training, fine-tuning, and evaluation pipelines for large-scale multimodal datasets. Define multi-task objectives and metrics that balance perception quality, geometric consistency, prediction accuracy, latency, and safety-critical behavior.\nInvestigate foundation-model approaches for autonomous driving, including vision-language models, multimodal pre-training, representation learning, and efficient deployment of learned world models.\nWork closely with perception, mapping, prediction, planning, simulation, data, and embedded-software teams to convert research advances into robust, production-quality AV systems.\nDevelop systematic analysis and debugging tools for model failures, cross-sensor disagreement, long-tail scenarios, distribution shift, and regressions in closed-loop simulation and on-road evaluation.

Qualifications

Minimum

BS, MS, or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a related technical field, or equivalent experience.\n8+ years of experience, with at least 2+ years in the AV or robotics industry and 2+ years of leadership experience in a technically area.\nStrong experience developing production-quality sensor-fusion, perception, state-estimation, or autonomous-driving systems.\nDemonstrated experience with learning-based multimodal perception or fusion involving two or more cameras, LiDAR, radar, map, navigation, and ego-motion signals.\nSolid understanding of 3D geometry, coordinate frames, calibration, temporal synchronization, ego-motion compensation, tracking, uncertainty estimation, and sensor failure modes.\nExperience with deep-learning methods for 3D perception, lego-context scene representation, occupancy/occlusion prediction, semantic segmentation, object detection/tracking, motion prediction, or planning.\nStrong C++ and Python programming skills, with hands-on experience developing, training, and optimizing deep-learning models in PyTorch. Experience with CUDA, distributed training, mixed-precision techniques, and efficient GPU inference using NVIDIA software and hardware is highly valued.\nExperience with Transformer, VLM, or multimodal foundation-model architectures, including pre-training, fine-tuning, distillation, quantization, or efficient inference.\nExperience training and evaluating models at scale, including distributed training, dataset curation, offline evaluation, simulation-based validation, and production monitoring.\nAbility to work across research and engineering boundaries: turn an ambiguous AV problem into measurable technical objectives, build the solution, and drive it to deployment.

Preferred

Experience building multi-task driving models that jointly predict perception, road structure, occupancy, motion, and/or trajectories from shared multimodal features.\nExperience with Transformer, VLM, or multimodal foundation-model architectures, including pre-training, fine-tuning, distillation, quantization, or efficient inference.\nExperience with BEV, point-cloud/voxel, neural scene representation, 3D reconstruction, occupancy-flow, or spatiotemporal world-model methods.\nPublications or open-source contributions in computer vision, robotics, machine learning, 3D perception, multimodal learning, or autonomous driving.\nExperience optimizing models for automotive-grade real-time deployment using NVIDIA GPUs, TensorRT, CUDA, or edge inference toolchains.