Researcher, Alignment Interpretability

OpenAI
San Francisco2026-09-01

About the job

OpenAI is seeking a researcher passionate about understanding deep networks, with a strong background in engineering, quantitative reasoning, and the research process. You will develop and carry out a research plan in mechanistic interpretability, in close collaboration with a highly motivated team. You will play a critical role in helping OpenAI ensure future models remain safe even as they grow in capability. This will make a significant impact on our goal of building and deploying safe AGI.

Responsibilities

- Develop and publish research on techniques for understanding representations of deep networks.

- Engineer infrastructure for studying model internals at scale.

- Collaborate across teams to work on projects that OpenAI is uniquely suited to pursue.

- Guide research directions toward demonstrable usefulness and/or long-term scalability.

Qualifications

Minimum

- Hold a Ph.D. or have research experience in computer science, machine learning, or a related field.

- Possess 2+ years of research engineering experience and proficiency in Python or similar languages.

Preferred

- Experience in the field of AI safety & alignment, mechanistic interpretability, or spiritually related disciplines.

- Enthusiasm for long-term AI safety & alignment and have thought deeply about technical paths to safe AGI.

- Thrive in environments involving large-scale AI systems and are excited to make use of OpenAI’s unique resources in this area.