NewXResearch hub
ExploreLibraryProfile
Account
Sign In
Henry Sleight
Scholar

Henry Sleight

Google Scholar ID: FRHn0z4AAAAJ
Research Manager, Anthropic Fellows Program, Program Manager, Constellation
AI SafetyAdversarial RobustnessModel Organisms of Misalignment
Homepage↗Google Scholar↗
Citations & Impact
All-time
Citations
395
 
H-index
9
 
i10-index
9
 
Publications
19
 
Co-authors
3
list available
Publications
16 items
AI Organizations are More Effective but Less Aligned than Individual Agents
2026
Cited
0
Abstractive Red-Teaming of Language Model Character
2026
Cited
1
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
2026
Cited
0
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
2025
Cited
0
Evaluating Control Protocols for Untrusted AI Agents
2025
Cited
0
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
2025
Cited
0
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
2025
Cited
0
Stress-Testing Model Specs Reveals Character Differences among Language Models
2025
Cited
0
Co-authors
3 total
Ethan Perez
Ethan Perez
Anthropic
John Hughes
John Hughes
Anthropic
Rylan Schaeffer
Rylan Schaeffer
Stanford University