About the job
NVIDIA is a global leader in high-speed computer vision, artificial intelligence (AI), and deep learning. Our team develops data engineering solutions that empower AI developers in autonomous vehicle (AV) domains to innovate quickly and effectively at scale. Are you ready to take on a senior technical role in building high-performance AI data pipelines? We seek an exceptional individual to design and optimize microservices and data pipelines to process massive volumes of AV data and enable seamless data mining and AI training. The ideal candidate will bring expertise in big data processing and distributed computing to create efficient solutions and overarching architectures for challenges such as video data curation, behavioral search, and AI dataset management.
Responsibilities
Scope and build tools, microservices, workflows, and distributed applications to accelerate data mining and AI training.
Design and implement solutions for streaming, resilience, logging, security, authentication, workflow orchestration, and data management.
Deploy AI models.
Design and develop Retrieval-Augmented Generation (RAG) workflows enabling hybrid and agentic patterns.
Analyze and operationalize complex distributed systems for speed-of-light performance.
Qualifications
Minimum
Experience developing high-performance, scalable software systems.
MS with 6+ years, or BS (or equivalent experience) with 8+ years of relevant experience in Computer Science, Computer Engineering, or a related technical field.
Strong programming skills in Python or Golang
Proficiency in key technologies like Kubernetes, Helm, Hive, Parquet, SQL, vector databases, e.g., Milvus.
Strong architectural skills with a proactive, problem-solving mentality.
Experience in data mining | AI development
Experience building ETL pipelines and working with big data engines.
Exceptional collaboration skills to work with system software and AI expert teams.
Eagerness to learn and adopt new technologies such as NVIDIA RAPIDS.
Preferred
Prior experience with large-scale real-time streaming, augmented reality, or data curation.
Prior background with Spark.
Exposure to the latest advances in AI, including Large Language Models, Vision-Language Models, and Retrieval-Augmented Generation (RAGs).
Innovative results, including patents, publications, or open source contributions.