Scholar
Sebastian Farquhar
Google Scholar ID: bvShhTEAAAAJ
Google DeepMind
AGI Alignment
Bayesian Deep Learning
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
5,760
H-index
26
i10-index
35
Publications
20
Co-authors
10
list available
Contact
No contact links provided.
Publications
7 items
GDM AI Control Roadmap
2026
Cited
0
Gram: Assessing sabotage propensities via automated alignment auditing
2026
Cited
0
Realistic honeypot evaluations for scheming propensity
2026
Cited
0
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
2026
Cited
0
An Approach to Technical AGI Safety and Security
2025
Cited
0
Do Multilingual LLMs Think In English?
2025
Cited
0
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
2025
Cited
0
Resume (English only)
Co-authors
10 total
Yarin Gal
Professor of Machine Learning, University of Oxford
Tom Rainforth
Associate Professor, University of Oxford
Jannik Kossen
FAIR, Meta
Lewis Smith
PhD Student, University of Oxford
Co-author 5
Michael A Osborne
Professor of Machine Learning, University of Oxford
Angelos Filos
Google DeepMind, University of Oxford
Tom Everitt
Staff Research Scientist at Google DeepMind