Scholar
Nina Panickssery
Google Scholar ID: 6-_i-jsAAAAJ
Anthropic
Language Models
AI Alignment
AI Interpretability
ML Safety
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
863
H-index
5
i10-index
4
Publications
7
Co-authors
12
list available
Publications
4 items
Mechanistically Eliciting Latent Behaviors in Language Models
2026
Cited
0
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
2025
Cited
0
Mitigating Many-Shot Jailbreaking
2025
Cited
0
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
arXiv.org · 2024
Cited
0
Co-authors
10 total
Meg Tong
Anthropic
Evan Hubinger
Member of Technical Staff, Anthropic
Julian Schulz
University of Göttingen
Andy Arditi
Northeastern University
Neel Nanda
Mechanistic Interpretability Team Lead, Google DeepMind
Wes Gurnee
Anthropic
Alexander Matt Turner
Research scientist, Google DeepMind
Daniel Paleka
ETH Zurich