NewXResearch hub
ExploreLibraryProfile
Account
Sign In
Euan Ong
Scholar

Euan Ong

Google Scholar ID: vT2qcI0AAAAJ
Anthropic
Machine LearningScience of Deep LearningMechanistic InterpretabilityAlgorithmic Reasoning
Homepage↗Google Scholar↗
Citations & Impact
All-time
Citations
334
 
H-index
7
 
i10-index
5
 
Publications
13
 
Co-authors
0
 
Publications
6 items
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
2026
Cited
0
Verbalizable Representations Form a Global Workspace in Language Models
2026
Cited
0
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
2025
Cited
0
Auditing language models for hidden objectives
2025
Cited
0
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
2025
Cited
0
Compact Proofs of Model Performance via Mechanistic Interpretability
arXiv.org · 2024
Cited
4