NewXResearch hub
ExploreLibraryProfile
Account
Sign In
Nina Panickssery
Scholar

Nina Panickssery

Google Scholar ID: 6-_i-jsAAAAJ
Anthropic
Language ModelsAI AlignmentAI InterpretabilityML Safety
Homepage↗Google Scholar↗
Citations & Impact
All-time
Citations
863
 
H-index
5
 
i10-index
4
 
Publications
7
 
Co-authors
12
list available
Publications
4 items
Mechanistically Eliciting Latent Behaviors in Language Models
2026
Cited
0
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
2025
Cited
0
Mitigating Many-Shot Jailbreaking
2025
Cited
0
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
arXiv.org · 2024
Cited
0
Co-authors
10 total
Meg Tong
Meg Tong
Anthropic
Evan Hubinger
Evan Hubinger
Member of Technical Staff, Anthropic
Julian Schulz
Julian Schulz
University of Göttingen
Andy Arditi
Andy Arditi
Northeastern University
Neel Nanda
Neel Nanda
Mechanistic Interpretability Team Lead, Google DeepMind
Wes Gurnee
Wes Gurnee
Anthropic
Alexander Matt Turner
Alexander Matt Turner
Research scientist, Google DeepMind
Daniel Paleka
Daniel Paleka
ETH Zurich