Applied Scientist / Machine Learning Engineer

Wayve
Sunnyvale2026-06-30

About the job

This role sits in the AI Platform organisation, on the data flywheel that powers every model we ship. The thesis is simple and compounding: the more intelligently we curate, enrich, and evaluate the real-world driving experience our fleet generates, the faster our foundation models improve, and the further they generalise across geographies, embodiments, and OEM platforms. As deployment scales, the bottleneck is shifting from raw model capacity to the quality and intelligence of the data engine and the rigour of how we measure progress. That is the problem you will own.

Responsibilities

Mine world-scale fleet data for rare, long-tail, and safety-critical events using active learning, smart sampling, and embedding-based retrieval and dedup.

Figure out what makes a good training dataset: which data, mix, and balance actually move the model, and turn that into repeatable curation across cities, sensor rigs, and embodiments.

Build high-quality enrichments that teams across the company depend on, through (semi-)automated enrichment and labeling pipelines and data quality at scale.

Build and fine-tune large-scale pretrained models, and run smaller-scale experiments to test and derisk ideas before committing serious compute.

Help build the best embodied VLM / VLA in the world for driving (the LINGO line): push multimodal perception, reasoning, language, and action.

Design rigorous offline and closed-loop evaluation: metrics and benchmarks that correlate with real on-road behaviour and safety, with deliberate coverage of rare and safety-critical scenarios.

Use world-model-based evaluation (GAIA) to probe counterfactual “what if” scenarios safely, repeatably, and at scale.

Contribute across the wider foundation-model stack as the work demands: generative world models (GAIA), policy learning, reinforcement learning, and reward modeling.

Qualifications

Minimum

A Masters with around 6 or more years of relevant experience, or a PhD with 2 or more years, in computer science, machine learning, robotics, mathematics, or a related field (required).

Strong ML and software fundamentals, and a track record of taking ML from research into production systems that run at scale.

Hands-on strength in one or more of: data curation, foundation model training, large-scale data wrangling, and foundation-model evaluation (for example, evaluation of LLMs or similar large models).

Experience with large-scale data and/or large neural networks, and the judgment to know which experiments and which data actually matter.

Fluency in Python and a modern deep-learning framework (PyTorch or similar), and comfort working with large, messy, real-world datasets.

Preferred

Autonomous driving, robotics, or other embodied-AI domains.

Foundation models, VLMs, world models, diffusion or autoregressive generative models, or reinforcement learning and reward modeling.

Large-scale data infrastructure: embedding and vector search (e.g. turbopuffer, Milvus), distributed data processing (Ray Data, Daft, Spark), lakehouse formats (Lance, Iceberg), or annotation tooling.

Closed-loop or simulation-based evaluation, and safety-critical ML.

Publications at top ML, CV, or robotics venues (NeurIPS, ICML, ICLR, CVPR, CoRL, RSS).