New research describes how to hold AI accountable for how well it protects users in crisis.
Content Warning: This article discusses suicide. For those experiencing suicidal thoughts or ideations, help is always available. The 988 Suicide & Crisis Lifeline provides free 24-hour, confidential support by calling or texting 988 or visiting 988lifeline.org.
We often judge chatbots by how they make us feel. Did the reply sound kind? Did it seem to understand? Did it avoid sounding cold or robotic? Those signals are important, but they are also not enough for serious mental health conversations, especially in the context of suicide risk. That project lies at the heart of “Benchmarking the Safety of General-Purpose Large Language Models for Suicide Risk Detection and Response,” a new working paper co-written by HBS AI Institute associate Julian De Freitas. To systematically address whether AI chatbots behave safely in critical situations, De Freitas and his co-authors examine how automated benchmarking is a new and important form of accountability: a way to make safety visible, comparable, and improvable.
Why This Matters
For business leaders and executives deploying AI in healthcare, HR, customer support, or any context involving users who may be at risk, this paper models what responsible benchmarking looks like: transparent criteria, standardized testing conditions, and honest acknowledgement of what the evaluation cannot yet capture. Just as a financial firm wouldn’t deploy an automated trading algorithm without rigorous stress-testing, true due diligence requires evaluating AI systems on their worst-case outputs, tracking specific failures, and building explicit, low-friction handoffs to human oversight when high-stakes thresholds are crossed.
Link to the HBS AI Institute Insight Article
Sign up for our newsletter to stay up to date with HBS AI Institute news and research