Scholar
Mantas Mazeika
Google Scholar ID: fGeEmLQAAAAJ
Center for AI Safety
ML Safety
AI Safety
Machine Ethics
ML Reliability
Follow
Google Scholar
↗
Citations & Impact
All-time
Citations
17,612
H-index
26
i10-index
28
Publications
20
Co-authors
5
list available
Publications
15 items
CheatBench: Measuring Reward Gaming in AI Agents
2026
Cited
0
Aggressive Compression Enables LLM Weight Theft
arXiv.org · 2026
Cited
0
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
2025
Cited
0
Remote Labor Index: Measuring AI Automation of Remote Work
2025
Cited
0
A Definition of AGI
2025
Cited
0
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
2025
Cited
0
TextQuests: How Good are LLMs at Text-Based Video Games?
2025
Cited
0
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
2025
Cited
0
Load more
Co-authors
5 total
Dan Hendrycks
Director of the Center for AI Safety (advisor for xAI and Scale)
Dawn Song
Professor of Computer Science, UC Berkeley
Andy Zou
PhD Student, Carnegie Mellon University
Bo Li
University of Illinois at Urbana–Champaign
David Forsyth
Professor of Computer Science, University of Illinois, Urbana Champaign