Scholar
Neel Nanda
Google Scholar ID: GLnX3MkAAAAJ
Mechanistic Interpretability Team Lead, Google DeepMind
AI
ML
AI Alignment
Interpretability
Mechanistic Interpretability
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
8,941
H-index
32
i10-index
45
Publications
20
Co-authors
9
list available
Contact
Email
neelnanda27@gmail.com
Twitter
Open ↗
Publications
40 items
Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
2026
Cited
0
How Transparent is DiffusionGemma?
2026
Cited
0
Subliminal Learning Is Steering Vector Distillation
2026
Cited
0
Building Better Activation Oracles
2026
Cited
0
How Well Do Models Follow Their Constitutions?
2026
Cited
0
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
2026
Cited
0
Simple LLM Baselines are Competitive for Model Diffing
2026
Cited
0
Emergent Misalignment is Easy, Narrow Misalignment is Hard
2026
Cited
0
Load more
Co-authors
8 total
Arthur Conmy
Google DeepMind
Senthooran Rajamanoharan
Google DeepMind
Wes Gurnee
Anthropic
Janos Kramar
Google DeepMind
Catherine Olsson
Anthropic
Lawrence Chan
PhD Student, UC Berkeley
Bilal Chughtai
Google DeepMind
Christopher Olah
Anthropic