About the job
We are hiring software engineers for the CUDA Tile team. NVIDIA GPUs are at the center of the deep learning revolution and continue to enable breakthroughs in generative AI, large language models, recommendation systems, speech recognition, image classification and other areas. Come join us to work with a top-notch team and have broad impact across the entire deep learning community.
Responsibilities
Design and implement compiler transformations for CUDA Tile, a new tile-based programming model for GPUs
Develop MLIR-based dialects and lowering passes
Optimize the performance of tile-based kernels to ensure efficient execution across multiple generations of NVIDIA GPU architectures
Define public APIs and craft compiler and optimization techniques
Perform general software engineering work including performance optimization
Qualifications
Minimum
Bachelors, Masters or Ph.D. in Computer Science, Computer Engineering or a related field (or equivalent experience)
3+ years of relevant work or research experience in compiler optimization, performance analysis and IR design
Ability to work independently, define project goals and scope, and lead your own development effort
Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design
Strong interpersonal skills along with the ability to work in a dynamic product-oriented team
Preferred
Knowledge of CPU and/or GPU architecture
CUDA or OpenCL programming experience
Experience with MLIR, LLVM, XLA, TVM and deep learning models and algorithms