New research describes how to hold AI accountable for how well it protects users in crisis.
Listen to this article:
Content Warning: This article discusses suicide. For those experiencing suicidal thoughts or ideations, help is always available. The 988 Suicide & Crisis Lifeline provides free 24-hour, confidential support by calling or texting 988 or visiting 988lifeline.org.
We often judge chatbots by how they make us feel. Did the reply sound kind? Did it seem to understand? Did it avoid sounding cold or robotic? Those signals are important, but they are also not enough for serious mental health conversations, especially in the context of suicide risk. That project lies at the heart of “Benchmarking the Safety of General-Purpose Large Language Models for Suicide Risk Detection and Response,” a new working paper co-written by HBS AI Institute associate Julian De Freitas. To systematically address whether AI chatbots behave safely in critical situations, De Freitas and his co-authors examine how automated benchmarking is a new and important form of accountability: a way to make safety visible, comparable, and improvable.
Key Insight: Defining the Yardstick
“To address the need for rigorous automated evaluation of AI chatbot safety, the Validation of Ethical and Responsible AI in Mental health (VERA-MH) framework was recently introduced.” [1]
For their study, the researchers leveraged the VERA-MH framework, an open-source tool built on an automated “LLM-as-a-judge” architecture. Rather than testing isolated one-off responses, VERA-MH simulated multi-turn conversations between an AI “user” persona (with hard-coded clinical characteristics like suicide risk level and healthcare access) and the chatbot being tested. A third model then acts as a judge, scoring the exchange against a structured rubric. This design is significant because risk can emerge, change, or be mishandled over time. A single empathetic response may look appropriate in isolation, but repeated vague reassurance, failure to ask follow-up questions, or poor guidance can become unsafe across a conversation. VERA-MH assesses five dimensions: detecting potential risk, confirming that risk through direct questioning, guiding users toward human care, maintaining a supportive tone, and staying within appropriate AI boundaries. The composite scoring formula is deliberately designed with a safety-first approach, penalizing harmful behavior more heavily than it rewards best practices.
Key Insight: Models are Better at Detecting Risk than Responding Safely
“[A]ll models still evidenced concerning rates of potentially harmful behavior, particularly around confirming risk and guiding to human care.” [2]
The results are both encouraging and sobering. Ten models across four providers (OpenAI, Anthropic, Google DeepMind, and xAI) were evaluated, and newer models improved over earlier ones for three of the four providers. But still, the highest overall score was 64 out of 100, and the researchers note that even a 65 could still reflect up to 20% of interactions rated as having high potential for harm. Nearly every model did well on detecting that a user might be at risk, but they failed to consistently act on it. In 60.8% of the analyzed conversations, chatbots failed to ask direct questions to confirm if an ambiguous user statement reflected explicit suicidal thoughts or another immediate safety risk. In about 33% of the conversations, models either failed to offer relevant mental healthcare resources or didn’t surface a specific 24/7 crisis line, and in another 12%, they failed to persistently direct users toward emergency support even when risk was immediate. Models also struggled with users who resisted recommendations and those lacking consistent healthcare access.
Key Insight: The Road Ahead
“[C]ontinued research and product work aimed to maximize ecological validity of personas is needed.” [3]
The paper is careful not to oversell benchmarking. VERA-MH is a tool for systematic evaluation, not a replacement for clinical trials, expert review, post-market monitoring, or real-world outcome studies. Additionally, the tests used simulated users rather than real people, models were accessed through APIs rather than consumer-facing interfaces, and very long conversations (where guardrails may degrade) were not fully evaluated.
Still, the authors conclude with several actionable iterations for the open-source benchmark and future model development. To prevent developers from simply “gaming” the open-source benchmark for higher scores without actual safety gains, independent third-party assessments remain vital. Future updates must introduce tougher adversarial scenarios, expand the base rate of “No Risk” personas to properly evaluate model specificity, and account for subtle, cumulative harms like AI sycophancy that unfold across multiple separate interactions.
Why This Matters
For business leaders and executives deploying AI in healthcare, HR, customer support, or any context involving users who may be at risk, this paper models what responsible benchmarking looks like: transparent criteria, standardized testing conditions, and honest acknowledgement of what the evaluation cannot yet capture. Just as a financial firm wouldn’t deploy an automated trading algorithm without rigorous stress-testing, true due diligence requires evaluating AI systems on their worst-case outputs, tracking specific failures, and building explicit, low-friction handoffs to human oversight when high-stakes thresholds are crossed.
Bonus
For another perspective on how AI systems designed to comfort users must be judged not by how caring they sound, but by whether they reliably protect users from foreseeable harm, check out Navigating the Promise and Peril of AI Companions for Older Adults.
References
[1] Bentley, Kate, et al., “Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response,” Harvard Business School Working Paper, No. 26-084 (May 2026): 3. https://dx.doi.org/10.2139/ssrn.6836551
[2] Bentley et al., “Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response,” 11.
[3] Bentley et al., “Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response,” 13.
Meet the Author

Julian De Freitas is an Assistant Professor of Business Administration in the Marketing Unit and Director of the Ethical Intelligence Lab at Harvard Business School, and Associate HBS AI Institute. His work sits at the nexus of AI, consumer psychology, and ethics.
Additional Authors: Kate H. Bentley, Emily Van Ark, Tim Hahn, Nicholas C. Jacobson, Nina Vasan, Ursula Whiteside, Luca Belli, Josh Gieringer, Nilu Zhao, Millard Brown, Adam M. Chekroud, Matt Hawrilenko
Watch a video version of the Insight Article here.