Deep Learning Compiler Engineer

Nvidia
US, CA, Santa Clara / US, TX, Austin / US, WA, Remote2026-08-19remote_local

About the job

We are hiring software engineers for the CUDA Tile team. NVIDIA GPUs are at the center of the deep learning revolution and continue to enable breakthroughs in generative AI, large language models, recommendation systems, speech recognition, image classification and other areas. Come join us to work with a top-notch team and have broad impact across the entire deep learning community.

Responsibilities

Design and implement compiler transformations for CUDA Tile, a new tile-based programming model for GPUs

Develop MLIR-based dialects and lowering passes

Optimize the performance of tile-based kernels to ensure efficient execution across multiple generations of NVIDIA GPU architectures

Define public APIs and craft compiler and optimization techniques

Perform general software engineering work including performance optimization

Qualifications

Minimum

Bachelors, Masters or Ph.D. in Computer Science, Computer Engineering or a related field (or equivalent experience)

3+ years of relevant work or research experience in compiler optimization, performance analysis and IR design

Ability to work independently, define project goals and scope, and lead your own development effort

Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design

Strong interpersonal skills along with the ability to work in a dynamic product-oriented team

Preferred

Knowledge of CPU and/or GPU architecture

CUDA or OpenCL programming experience

Experience with MLIR, LLVM, XLA, TVM and deep learning models and algorithms