π€ AI Summary
Current safety evaluations of mental healthβfocused AI systems predominantly rely on small-scale simulated benchmarks, which inadequately capture the linguistic and contextual diversity of real-world scenarios. This study presents the first systematic safety assessment combining replication across four established benchmarks with an ecological audit of 20,000 real user conversations, comparing specialized mental health AI against six state-of-the-art general-purpose large language models on high-risk topics. Employing clinical expert blind review, LLM-based adjudicators, automated crisis resource triggering, and statistical confidence interval analysis, the findings reveal that the specialized system exhibits significantly lower rates of harmful content in response to prompts involving self-harm, eating disorders, and substance abuse. In live deployment, it achieved zero end-to-end missed detections, with the LLM adjudicator demonstrating 100% sensitivity and 99.2% specificity. The work advocates for ecological auditing as a critical complement to pre-deployment safety testing.
π Abstract
Large language models (LLMs) are increasingly used for mental health, yet safety evaluations rely primarily on small, simulation-based benchmarks removed from real-world language. We replicate four published safety evaluations assessing suicide risk handling, harmful content generation, and jailbreak resistance for general-purpose frontier models and a purpose-built mental health AI. We then conduct an ecological audit of 20,000 real user conversations with the purpose-built system, which includes layered safeguards for suicide and non-suicidal self-injury (NSSI). The purpose-built AI was significantly less likely than general-purpose LLMs to produce harmful content across suicide/NSSI (.4β11.27% vs 29.0β54.4%), eating disorder (8.4% vs 54.0%), and substance use (9.9% vs 45.0%) benchmarks. In real user data, clinician review found zero suicide-risk cases without crisis resources. Three NSSI mentions (.015%) lacked intervention, implying a .38% lower-bound false negative rate. Findings support the utility of ecological audits for safety estimation.