NewXResearch hub
ExploreLibraryProfile
Account
Sign In
Neel Nanda
Scholar

Neel Nanda

Google Scholar ID: GLnX3MkAAAAJ
Mechanistic Interpretability Team Lead, Google DeepMind
AIMLAI AlignmentInterpretabilityMechanistic Interpretability
Homepage↗Google Scholar↗
Citations & Impact
All-time
Citations
8,941
 
H-index
32
 
i10-index
45
 
Publications
20
 
Co-authors
9
list available
Contact
Emailneelnanda27@gmail.comTwitterOpen ↗
Publications
40 items
Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
2026
Cited
0
How Transparent is DiffusionGemma?
2026
Cited
0
Subliminal Learning Is Steering Vector Distillation
2026
Cited
0
Building Better Activation Oracles
2026
Cited
0
How Well Do Models Follow Their Constitutions?
2026
Cited
0
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
2026
Cited
0
Simple LLM Baselines are Competitive for Model Diffing
2026
Cited
0
Emergent Misalignment is Easy, Narrow Misalignment is Hard
2026
Cited
0
Co-authors
8 total
Arthur Conmy
Arthur Conmy
Google DeepMind
Senthooran Rajamanoharan
Senthooran Rajamanoharan
Google DeepMind
Wes Gurnee
Wes Gurnee
Anthropic
Janos Kramar
Janos Kramar
Google DeepMind
Catherine Olsson
Catherine Olsson
Anthropic
Lawrence Chan
Lawrence Chan
PhD Student, UC Berkeley
Bilal Chughtai
Bilal Chughtai
Google DeepMind
Christopher Olah
Christopher Olah
Anthropic