Scholar
Arthur Conmy
Google Scholar ID: n4HIyXQAAAAJ
Google DeepMind
AGI Safety
AI Safety
Interpretability
Mechanistic Interpretability
Machine Learning
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
2,569
H-index
18
i10-index
20
Publications
20
Co-authors
20
list available
Publications
19 items
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
2026
Cited
0
How Transparent is DiffusionGemma?
2026
Cited
0
Subliminal Learning Is Steering Vector Distillation
2026
Cited
0
How do LLMs Compute Verbal Confidence
2026
Cited
0
Automatically Finding Reward Model Biases
2026
Cited
0
Simple LLM Baselines are Competitive for Model Diffing
2026
Cited
0
Fluid Representations in Reasoning Models
2026
Cited
0
Building Production-Ready Probes For Gemini
2026
Cited
1
Load more
Co-authors
16 total
Neel Nanda
Mechanistic Interpretability Team Lead, Google DeepMind
Senthooran Rajamanoharan
Google DeepMind
Janos Kramar
Google DeepMind
Aengus Lynch
University College London
Rohin Shah
Research Scientist, Google DeepMind
Jacob Steinhardt
Stanford University
Rowan Wang
Unknown affiliation
Iván Arcuschin
Independent Researcher