Senior+ Software Engineer - Research Platform, Consumer Devices

OpenAI
San Francisco, CA, USA2026-06-26Hybrid

About the job

We are looking for a Software Engineer to join our team to build tools and services that enable AI research, evaluation, and data generation workflows. The best work in this role will start with an ambiguous design question and turn it into working research systems. You will work closely with researchers, designers, and engineers to build the evaluation systems, synthetic data generation pipelines, review tools, and supporting platform services.

Responsibilities

- Build web applications, APIs, data models, and backend services for AI research workflows.

- Build tools to author and manage evaluation tasks, rubrics, graders, suites, and rollout configurations, including workflows for publishing, versioning, auditing, and sharing research artifacts.

- Automate evaluation runs and generate useful reports for design, research, and engineering teams.

- Support synthetic data generation workflows for multimodal and conversational research, including tools that combine transcripts, media, and model comparisons.

- Translate product and research questions into measurable scenarios, automated graders, and human-evaluation campaigns, and develop measures of task quality, coverage, diversity, and semantic spread.

- Diagnose issues across application code, workers, model endpoints, deployments, and compute infrastructure, and improve reliability through health checks, observability, reproducible launch paths, data integrity safeguards, and automated verification.

- Lead migrations and dependent changes across research tools, evaluation systems, and supporting services.

- Partner closely with designers, model researchers, research engineers, and infrastructure teams, and onboard contributors to create high-quality evaluation and synthetic-data workflows.

Qualifications

Minimum

- Have 7+ years of professional software engineering experience.

- Have strong full-stack experience across web applications, backend services, APIs, and data models, including ownership of complex systems spanning multiple services or repositories.

- Have expertise in generative AI, multimodal models, or model-evaluation systems.

- Have built effective internal tools for both technical and non-technical users.

- Are comfortable debugging distributed workflows and production infrastructure.

- Have strong product judgment and can translate ambiguous requirements into concrete plans.

- Are energized by working between designers and researchers in a multidisciplinary team, connecting qualitative judgment to rigorous evidence.

- Communicate clearly and work effectively across engineering, design, and research.

Preferred

- Expertise in synthetic data generation, simulation, conversational AI, speech, video, motion, or embodied interaction.

- Experience with automated graders, human evaluation, supervised fine-tuning, reinforcement learning, or experiment- and dataset-management platforms.

- Operated GPU-backed inference or rollout workloads at very large scale.