Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

πŸ“… 2026-01-14
πŸ›οΈ Research Square
πŸ“ˆ Citations: 7
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Current safety evaluations of mental health–focused AI systems predominantly rely on small-scale simulated benchmarks, which inadequately capture the linguistic and contextual diversity of real-world scenarios. This study presents the first systematic safety assessment combining replication across four established benchmarks with an ecological audit of 20,000 real user conversations, comparing specialized mental health AI against six state-of-the-art general-purpose large language models on high-risk topics. Employing clinical expert blind review, LLM-based adjudicators, automated crisis resource triggering, and statistical confidence interval analysis, the findings reveal that the specialized system exhibits significantly lower rates of harmful content in response to prompts involving self-harm, eating disorders, and substance abuse. In live deployment, it achieved zero end-to-end missed detections, with the LLM adjudicator demonstrating 100% sensitivity and 99.2% specificity. The work advocates for ecological auditing as a critical complement to pre-deployment safety testing.
πŸ“ Abstract
Large language models (LLMs) are increasingly used for mental health, yet safety evaluations rely primarily on small, simulation-based benchmarks removed from real-world language. We replicate four published safety evaluations assessing suicide risk handling, harmful content generation, and jailbreak resistance for general-purpose frontier models and a purpose-built mental health AI. We then conduct an ecological audit of 20,000 real user conversations with the purpose-built system, which includes layered safeguards for suicide and non-suicidal self-injury (NSSI). The purpose-built AI was significantly less likely than general-purpose LLMs to produce harmful content across suicide/NSSI (.4–11.27% vs 29.0–54.4%), eating disorder (8.4% vs 54.0%), and substance use (9.9% vs 45.0%) benchmarks. In real user data, clinician review found zero suicide-risk cases without crisis resources. Three NSSI mentions (.015%) lacked intervention, implying a .38% lower-bound false negative rate. Findings support the utility of ecological audits for safety estimation.
Problem

Research questions and friction points this paper is trying to address.

mental-health AI safety
ecological auditing
simulation-based benchmarks
real-world conversations
harmful content
Innovation

Methods, ideas, or system contributions that make the work stand out.

ecological auditing
mental-health AI safety
real-world deployment evaluation
clinician-adjudicated validation
harmful content mitigation
πŸ”Ž Similar Papers
No similar papers found.
C
Caitlin A. Stamatis
Slingshot AI, New York, NY, USA
J
Jonah Meyerhoff
Northwestern University Feinberg School of Medicine, Chicago, IL, USA
Richard Zhang
Richard Zhang
Senior Research Scientist, Adobe
Computer VisionMachine LearningDeep LearningComputer Graphics
O
Olivier Tieleman
Slingshot AI, New York, NY, USA
M
Matteo Malgaroli
New York University School of Medicine, New York, NY, USA
T
Thomas D. Hull
Slingshot AI, New York, NY, USA