About the job
We aim to turn production data into intelligence on the safety of deployed AI models. Safety Oversight is a new team tasked with using large-scale production traffic and a variety of automated evaluation methods to monitor the safety and alignment of deployed models. Our work will ensure we measure the real-world efficacy of our safety stack—both of in-model safety training and out-of-model safety mitigations to ensure we are effective in our goal of deploying safe models that are used for widespread public benefit.
Responsibilities
Build classifiers and large-scale data pipelines to detect model misbehavior and misuse end-to-end.
Research and develop cross-context monitoring systems to detect coordinated harms, developing novel signal aggregation methods across disparate user sessions to identify large-scale attack vectors.
Think critically about novel methods for monitoring using model activations, actions, chains-of-thought and final answers.
Collaborate closely with infrastructure teams and data scientists to scale your work and regularly share results with the wider safety team.
Qualifications
Minimum
PhD in Computer Science, a related field, or equivalent practical experience.
3 years of experience building and shipping technical products.
Experience in the domain area of generative AI and Large Language Models (LLM).
Preferred
Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
3 years of experience developing code, running experiments and analyses collaboratively with coding agents.
Experience building large-scale, highly parallelised data pipelines, working on data quality, automated evaluation design and simple statistical modeling.
Ability to approach new research questions and implement technical solutions for them at scale.
Ability to use AI every day to build and find ways to push the frontier of model capabilities to accelerate work.