NewXResearch hub
ExploreLibraryProfile
Account
Sign In
Andy Arditi
Scholar

Andy Arditi

Google Scholar ID: NgyIgX4AAAAJ
Northeastern University
Interpretability
Homepage↗Google Scholar↗
Citations & Impact
All-time
Citations
367
 
H-index
6
 
i10-index
5
 
Publications
10
 
Co-authors
7
list available
Contact
Emailandyrdt@gmail.comTwitterOpen ↗
Publications
7 items
Learning the identity: a case study of how SGD selects among functional decompositions
2026
Cited
0
Synthetic Persona Pretraining: Alignment from Token Zero
2026
Cited
0
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
2025
Cited
0
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
2025
Cited
0
Inverse Scaling in Test-Time Compute
2025
Cited
0
Adversarial Manipulation of Reasoning Models using Internal Representations
2025
Cited
0
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
2025
Cited
0
Resume
Background
  • Research Interest: AI interpretability
Miscellany
  • Contact Information: andyrdt@gmail.com, andyarditi
Co-authors
6 total
Neel Nanda
Neel Nanda
Mechanistic Interpretability Team Lead, Google DeepMind
Wes Gurnee
Wes Gurnee
Anthropic
Nina Panickssery
Nina Panickssery
Anthropic
Daniel Paleka
Daniel Paleka
ETH Zurich
Runjin Chen
Runjin Chen
PHD student at UT Austin
Jack Lindsey
Jack Lindsey
Anthropic