About the job
At Nvidia we meet customers where they are on their AI journey on our GPUs - this means we build best in class frameworks in OSS and support a robust eco system of other OSS frameworks. and As a Product Manager for AI Platform post-training and RL, you will be responsible for building the tools, SDKs, and libraries which enable model builders to get the best large scale performance, resilience and experience on NVIDIA GPUs.
Responsibilities
Create and optimize post-training RL libraries, frameworks to help researchers and production model builders
Develop product strategy, roadmaps, and go-to-market plans
Collaborate with internal and external customers to build product-based roadmaps for training/post training software
Work with leadership to align with and drive company strategy
Qualifications
Minimum
Experience with design and scaling of training/post training and optimization software (ex. VeRL, Tunix, Nemo Framework, PyTorch distributed, torchtitan, etc.)
Demonstrable knowledge of GenAI or machine learning concepts, particularly around model training, performance optimization, and software development and delivery
Experience with large scale distributed systems
BS or MS degree in Computer Science, Computer Engineering, or similar experience (or equivalent experience)
15+ years of technical product management, or similar, experience at a technology company
Strong communication and interpersonal skills
Preferred
Experience leading RL for research to production at scale
Working on Open Source & Github-first developer products with deep customer interactions
Knowledge of GPU architecture, HW/SW co-design, and performance profiling