About the job
We are now looking for a Senior DFX Software Engineer - Machine Learning. Do you like to think creatively and enjoy solving challenges that require innovation? If so, we may have an opportunity for you. On our team we define and build methodologies, software, and flows tailored to the field of silicon device testing, silicon debug, and silicon failure analysis. We owe our success to our people, some of the brightest in the world, and a company culture that fosters and encourages innovation, and fuels our creativity!
Our team contributes to the advancement of all the fields in which NVIDIA participates, from gaming to building groundbreaking state-of-the-art compute platforms and Artificial Intelligence, by enabling high quality Silicon defect screening to sustain all these fields. This often requires new ways of thinking in order to meet new challenges, and we pride ourselves in our ability to tackle these challenges in ways that enable our success. If you are a like-minded person who enjoys innovation and likes to solve technical challenges then we would love to hear from you!
Responsibilities
Develop high performance software to enable design and development of efficient test pattern generation, application of these patterns on Silicon, failure analysis, and yield learning
Create efficient parallel graph traversal and graph analysis techniques
Work with multi-functional teams to assess and tackle problems that involve multiple areas of expertise through the company
Apply LLMs, RAGs, graph-based ML approaches, and reinforcement learning to define innovative solutions
Qualifications
Minimum
BS in EE or CS (or equivalent experience)
5+ years of experience in Software development
Strong programming experience in Python or C++
Experience use of LLMs (Large Language Models), GNNs (Graph Neural Networks), and Reinforcement Learning for efficient EDA solution
Expertise in high performance algorithms for DFT, simulations, and failure analysis
Understanding of different agent architectures, RAG systems, and communication protocols
Deep familiarity with reinforcement learning algorithms like PPO, SAC, or Q-learning, including experience tuning hyper-parameters and reward functions
Hands on experience with large scale training (e.g., ZeRO) and data processing (e.g. Spark)
Excellent communication skills
Preferred
MS or higher degree preferred
Hands on development in modern C++ is a huge plus
Proven deployment of large-scale agentic application with high concurrency and agility
Experience with software and hardware especially involving DFT, failure analysis, and CAD tools
Working experience of agentic models / frameworks, observability and evaluation tools
In-depth understanding of the graph neural networks, and reinforcement learning for logic design automation
Experience with fine-tuning large language models, building advanced multi-agent systems, RAG pipelines and vector databases