Scholar
Nora Belrose
Google Scholar ID: p_oBc64AAAAJ
Research Lead, EleutherAI
interpretability
neural networks
transformers
nlp
ai
Follow
Google Scholar
↗
Citations & Impact
All-time
Citations
703
H-index
10
i10-index
10
Publications
20
Co-authors
14
list available
Publications
16 items
The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
2026
Cited
0
Bergson: An Open Source Library for Data Attribution
2026
Cited
0
Binary Sparse Coding for Interpretability
2025
Cited
0
Evaluating SAE interpretability without explanations
2025
Cited
0
Mechanistic Anomaly Detection for"Quirky"Language Models
2025
Cited
0
Slowing Learning by Erasing Simple Features
2025
Cited
0
Examining Two Hop Reasoning Through Information Content Scaling
2025
Cited
0
Converting MLPs into Polynomials in Closed Form
2025
Cited
0
Load more
Co-authors
11 total
Stella Biderman
EleutherAI
Adam Gleave
CEO at FAR AI
Zach Furman
Research Fellow, Timaeus
Lev McKinney
University of Toronto
Jacob Steinhardt
Stanford University
Sergey Levine
UC Berkeley, Physical Intelligence
Michael Dennis
Google DeepMind
Tom Tseng
FAR AI