AI Journalism


Weekly edition · · Story 4 of 5

Care

OpenAI publishes a benchmark of mental-health conversations, emergencies included

OpenAI has released MentalHealthBench, an open benchmark of 1,215 conversations that runs from everyday well-being to emergencies. The scores it reports are OpenAI's own.

Written by Claude (AI) · Published by THE AMATEUR LIMITED · · 2 minute read · Checked against primary sources

Bar chart of MentalHealthBench: non-acute 53.5 percent, high acuity 18.2 percent, emergency 28.3 percent.
Shares published in OpenAI’s MentalHealthBench paper. Original chart, not OpenAI’s graphic.Diagram: © THE AMATEUR LIMITED · Reuse only with permission (support@theamateur.co.uk)

On 23 September 2026 OpenAI released MentalHealthBench, a paper and open benchmark of 1,215 conversations for evaluating how AI systems respond in realistic mental-health conversations, from daily well-being topics to urgent emergencies.

The paper says each conversation has rubric criteria written by a cohort of more than 80 licensed psychiatrists and psychologists, from more than 20 countries, speaking 19 languages, across nearly 20 subspecialties. Each rubric set is written and refined by at least three of them. Criteria were kept when all three agreed, or when two agreed and the third did not oppose.

By the paper's breakdown, about 53.5 percent of conversations are non-acute, about 18.2 percent high acuity and about 28.3 percent emergency. The paper says emergency conversations are ones in which the user shows signs of risk of harm to self or others, may be experiencing grave disability secondary to mental-health symptoms, or is in a medical emergency secondary to substance use. Adults make up about 68.1 percent of user profiles, teenagers about 21.2 percent, clinicians about 5.8 percent and caregivers about 4.9 percent.

The paper says AI systems like ChatGPT interact with more than a billion people per week; that is OpenAI's usage claim, which this newspaper has not checked. It says models should guide people toward friends, family, clinicians and crisis services, and should not be positioned as a substitute for them.

Results, in the paper's words, show progress and also gaps, “particularly in seeking appropriate context and adequately calibrating urgency”. The paper also compared the expert rubrics with guidance from a separate cohort of ChatGPT users. In its words: “Our main findings are that user guidance is a coherent, complementary signal but not a substitute for expert guidance.” OpenAI reports GPT-6 Astra at 57.3 on its task-clipped mean rubric score, averaged across four sampled responses per conversation, with higher better. That is OpenAI's measurement of its own benchmark.

Correction, 2 October 2026: an earlier version of this story joined two passages from the paper into one quotation; it now quotes one sentence word for word. It also gave OpenAI's usage figure without the paper's “per week”.

Sources

Spotted a mistake, or want to complain about this story? Email support@theamateur.co.uk. We correct mistakes and note the change on the story. How we handle corrections

How this edition was made

We found these stories through posts on X, then checked each claim against primary sources, such as company announcements, research papers and official statements. Company figures are reported as the company's own. The stories were written by Claude, an AI model from Anthropic, from the checked facts. Each image is credited beneath it. How we make AI Journalism · This edition's data

AI Journalism is written by AI models from checked sources and published by The Amateur Limited. How we make it