About the job
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation.
Responsibilities
Develop scalable and sustainable system architecture and designs for products, services and enhancements.">">Defend performance of critical services in alignment with customer expectations and SLOs.">">Own and define strategy and set direction and establish roadmaps for Vertex AI services to increase reliability, efficiency and ultimately feature velocity.">">Resolve outages or service disruptions and help design solutions to ensure systems are protected from similar classes of problems in the future.">">Collaborate with development counterparts to incorporate and deliver enhancements to systems resulting in improved reliability, scalability and or performance.
Qualifications
Minimum
Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.">">8 years of experience with software development in one or more programming languages.">">2 years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering roles managing large-scale infrastructure.">">3 years of experience designing, analyzing, and troubleshooting distributed systems.">">3 years of experience in machine learning infrastructure.
Preferred
Master's degree in Computer Science or Engineering.">">Experience enhancing and supporting large production systems on compute infrastructure.">">Experience with networking, capacity and performance.">">Experience in large-scale systems, architecture design and complex system integrations or migrations.">">Experience supporting a Tier 1 rotation.">">Expertise in SRE production principles and best practices.