Principal Platform Software Engineer

Oracle
US2026-08-14

About the job

Oracle Cloud Infrastructure’s (OCI) Developer Platform Builder Tools organization builds the next generation of developer productivity platforms, intelligent development workflows, and AI-powered engineering tools that accelerate software delivery across the enterprise.

We are seeking a Principal Platform Software Engineer to define the architecture and technical direction of scalable services that improve the developer experience and increase engineering productivity. You will work across platform engineering, distributed systems, cloud-native technologies, machine learning, large language models (LLMs), and developer tooling.

A key focus of this role is building a new AI-native testing platform that learns real-world service behavior, generates realistic test traffic, safely validates workloads, and detects regressions with minimal manual test authoring. Its closed-loop workflow observes changing traffic patterns, executes and evaluates tests, and continuously improves future coverage.

As a principal engineer, you will shape the platform from foundational architecture through production deployment and broad adoption across OCI. You will solve ambiguous, organization-wide technical problems, establish engineering standards, and provide technical leadership to senior engineers, product managers, data scientists, and OCI service teams.

Responsibilities

Define the technical vision, architecture, and multiyear evolution of OCI’s AI-native developer platforms.

Lead the design and delivery of secure, highly available services that operate reliably across large-scale, distributed cloud environments.

Architect systems that analyze service behavior, model traffic patterns, generate realistic workloads, and evaluate functional and non-functional outcomes.

Establish agent-assisted workflows that plan, execute, and evaluate canary, functional, integration, load, and performance tests within clearly defined safety controls.

Guide the application of machine learning and LLM capabilities to service telemetry, API changes, incidents, test results, and engineering knowledge.

Develop architectural approaches for failure diagnosis, regression detection, workload modeling, and the conversion of production incidents into reusable test coverage.

Define platform APIs, data models, extension points, and integration patterns that enable adoption across diverse OCI services and engineering environments.

Set engineering standards for scalability, reliability, observability, security, privacy, performance, and responsible AI.

Define success metrics and feedback loops that measure test effectiveness, traffic-pattern coverage, developer effort saved, platform reliability, and adoption.

Provide hands-on technical leadership, mentor senior engineers, resolve cross-team architectural challenges, and raise engineering standards across the organization.

Qualifications

Minimum

Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.

Eight or more years of professional software engineering experience, including significant experience designing and operating large-scale production systems.

Proficiency in one or more modern programming languages, such as Java, Go, Python, or C++.

Deep expertise in distributed systems, system design, APIs, concurrency, data modeling, and scalable service architectures.

Demonstrated success defining architecture and delivering complex platforms, backend services, microservices, event-driven systems, or large-scale data pipelines.

Experience developing and operating cloud-native software using CI/CD, containerization, orchestration, and infrastructure automation.

Strong understanding of service reliability, observability, performance engineering, security, and production operations.

Experience integrating machine learning, generative AI, or other intelligent capabilities into production software systems.

Proven ability to lead high-impact technical initiatives across multiple teams, influence without direct authority, and navigate ambiguous requirements.

Excellent technical judgment and communication skills, including the ability to explain complex architectural decisions to engineering, product, and executive stakeholders.

A track record of mentoring senior engineers and improving engineering practices across teams or organizations.

Preferred

Experience architecting developer platforms, CI/CD systems, testing infrastructure, or engineering productivity tools at enterprise scale.

Experience developing production solutions using machine learning, generative AI, or LLMs.

Familiarity with AI agents, tool-using models, retrieval-augmented generation, prompt engineering, LLM-as-a-judge approaches, or model evaluation.

Experience with machine learning frameworks such as PyTorch, TensorFlow, Hugging Face Transformers, or equivalent technologies.

Experience with traffic replay, workload modeling, canary analysis, load testing, performance testing, or chaos engineering.

Deep experience with Kubernetes, containers, service meshes, and OCI or another major public cloud.

Experience designing extensible, multi-tenant platforms used by multiple engineering organizations.

Experience analyzing API schemas, code changes, service dependencies, production incidents, and traffic patterns.

Familiarity with AI safety controls, model monitoring, privacy, and secure enterprise data handling.