25 frontier models spot users in crisis but the median one fails to hand off to help in ~35% of cases
Confirmed
The takeaway
On Oct 9, 2026 Scale AI published DistressBench (shared first with TIME): 19 licensed clinicians and crisis counselors wrote 718 realistic crisis conversations, and 25 frontier models from OpenAI, Anthropic, Google and others were scored.
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 3 of 5
- Last verified
- 10 October 2026
Your AI and this story
- GPT-6 Astra162 days after its cutoff
- Claude Opus 5.5101 days after its cutoff
- Gemini 3.8 Flash192 days after its cutoff
- Grok 4.7131 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 101 days before it.
Key facts
- Authors: Patrick Oathout (Scale AI red team and safety lead) and Drew Rein; published Oct 9, 2026
- Data: 718 chats written by 19 licensed clinicians and crisis counselors; 25 frontier models tested
- Rubric: compassion, de-escalation, expert referral, non-moralizing responses, therapist disclaimer
- Weighted score: best model ~88%, median ~71%, clinician reference ~99%
- Missed crisis handoff in recognised crises: median model ~35%, best ~16%. Median de-escalation rate ~46% (best ~75%)
- Models picked the better reply 96% of the time when asked to judge, yet failed to give it ~85% of the time (Scale blog)
- Multi-turn conversations scored ~69% vs ~74% single-turn; in extended adversarial sessions 18 of 20 conversations ended in documented violations
- Scale’s blog post does not name individual models’ scores in its main text
What happened
Scale AI built a clinician-written benchmark of suicidal-ideation and self-harm conversations and scored 25 frontier chatbots on how they respond. The main gap is in action, not detection: models notice distress but often do not point the user to human help, and long conversations make it worse.
Why it matters
Chatbot behaviour in crises is at the centre of 2026 lawsuits and state laws on companion bots. A clinician-authored, multi-turn benchmark gives a measurable target, and its result (18 of 20 long adversarial sessions ending in violations) suggests that single-turn safety tests overstate real-world safety.
Sources
3 sources from 2 sites. Numbers match the chips in the text.
3 sources: 2 primary, 1 press
Primary
- Scale AI: DistressBench (Oct 9, 2026)labs.scale.com, official
- Scale AI: DistressBench paperlabs.scale.com, paper
Press
Changes
- Filed from Scale AI’s blog post and TIME’s exclusive