OpenAI releases MentalHealthBench, an open benchmark for AI in mental-health conversations
On Sept 23, 2026 OpenAI released MentalHealthBench: 1,215 synthetic mental-health conversations with 5,262 rubric criteria written with 80+ licensed clinicians from 22 countries. It scores safety, context-seeking, user agency and actionable guidance. Reported top scores: GPT-6 Astra 57.3%, GPT-6 Sol 53.9%, Claude Opus 5.5 52.4%, and GPT-4o 32.1%.
Key facts
- 1,215 conversations, 5,262 rubric criteria, 80+ mental-health experts from 22 countries
- Scenario mix: 53.5% non-acute, 18.2% high-acuity, 28.3% emergency
- 10 behavioral axes, incl. context seeking, empathy, urgency calibration and reality testing
- Reported scores: GPT-6 Astra 57.3%, GPT-6 Sol 53.9%, Claude Opus 5.5 52.4%, GPT-6 Luna 50.2%, Muse Spark 1.3 47%, GPT-4o 32.1%, Gemini 2.5 Pro 29.5%
- Critics (e.g. NxCode) note that rubric scoring of single conversations cannot measure long-run outcomes for users
What happened
OpenAI published an expert-written, rubric-graded benchmark for mental-health conversations, extending its HealthBench approach. It was released as an open benchmark, and the launch results compare OpenAI models with Claude, Gemini and Meta's Muse Spark.
Why it matters
Mental-health use of chatbots was a major 2025–2026 safety and litigation topic, and this gives labs and regulators a shared measure. The scores are OpenAI-reported and, per secondary sources, OpenAI models lead. Treat them as vendor results (confidence: medium, since openai.com was not directly readable).
Changelog
- 2026-09-29: created
Related events
- OpenAI releases GPT-6 Astra, its first GPT-6 model ★★★★★
- Anthropic releases Claude Opus 5.5 — Fable-5.1-level performance at $4/$20, first model of the Claude 5.5 family ★★★★★
Sources (4)
- officialOpenAI: Introducing MentalHealthBench
- pressAI Weekly: OpenAI releases MentalHealthBench with 1,215 conversations from 80+ psychologists
- pressEdTech Innovation Hub: OpenAI launches MentalHealthBench
- discussionNxCode: MentalHealthBench can score an AI's advice. It cannot tell…
id: 2026-09-23-openai-mentalhealthbench · updated 2026-09-29 · open in the interactive timeline