{
  "id": "mental-health",
  "edition": "2026-09-23",
  "position": 4,
  "role": "secondary",
  "kind": "story",
  "kicker": "Care",
  "headline": "OpenAI publishes a benchmark of mental-health conversations, emergencies included",
  "standfirst": "OpenAI has released MentalHealthBench, an open benchmark of 1,215 conversations that runs from everyday well-being to emergencies. The scores it reports are OpenAI's own.",
  "body": [
    "On 23 September 2026 OpenAI released MentalHealthBench, a paper and open benchmark of 1,215 conversations for evaluating how AI systems respond in realistic mental-health conversations, from daily well-being topics to urgent emergencies.",
    "The paper says each conversation has rubric criteria written by a cohort of more than 80 licensed psychiatrists and psychologists, from more than 20 countries, speaking 19 languages, across nearly 20 subspecialties. Each rubric set is written and refined by at least three of them. Criteria were kept when all three agreed, or when two agreed and the third did not oppose.",
    "By the paper's breakdown, about 53.5 percent of conversations are non-acute, about 18.2 percent high acuity and about 28.3 percent emergency. The paper says emergency conversations are ones in which the user shows signs of risk of harm to self or others, may be experiencing grave disability secondary to mental-health symptoms, or is in a medical emergency secondary to substance use. Adults make up about 68.1 percent of user profiles, teenagers about 21.2 percent, clinicians about 5.8 percent and caregivers about 4.9 percent.",
    "The paper says AI systems like ChatGPT interact with more than a billion people per week; that is OpenAI's usage claim, which this newspaper has not checked. It says models should guide people toward friends, family, clinicians and crisis services, and should not be positioned as a substitute for them.",
    "Results, in the paper's words, show progress and also gaps, “particularly in seeking appropriate context and adequately calibrating urgency”. The paper also compared the expert rubrics with guidance from a separate cohort of ChatGPT users. In its words: “Our main findings are that user guidance is a coherent, complementary signal but not a substitute for expert guidance.” OpenAI reports GPT-6 Astra at 57.3 on its task-clipped mean rubric score, averaged across four sampled responses per conversation, with higher better. That is OpenAI's measurement of its own benchmark."
  ],
  "word_count": 314,
  "reading_minutes": 2,
  "url": "https://theamateur.co.uk/ai-journalism/2026-09-23/mental-health/",
  "api_url": "https://theamateur.co.uk/ai-journalism/api/v1/stories/2026-09-23/mental-health.json",
  "published_at": "2026-09-23T17:00:00+01:00",
  "updated_at": "2026-09-23T17:00:00+01:00",
  "sources": [
    {
      "role": "primary",
      "label": "MentalHealthBench paper",
      "url": "https://cdn.openai.com/ctf-cdn/MentalHealthBench_A_Comprehensive_Benchmark_of_AI_Capabilities_in_Realistic_Mental_Health_Conversations.pdf"
    },
    {
      "role": "primary",
      "label": "OpenAI, 23 September 2026",
      "url": "https://openai.com/index/introducing-mentalhealthbench/"
    },
    {
      "role": "independent",
      "label": "Washington Post, via Newsday",
      "url": "https://www.newsday.com/business/chatbots-mental-health-ai-yg8dl7wr"
    },
    {
      "role": "discovery",
      "label": "OpenAI on X",
      "url": "https://x.com/OpenAI/status/2102837574092161102"
    }
  ],
  "images": [
    {
      "id": "mental-health-hero",
      "role": "hero",
      "editorial_role": "mechanism",
      "url": "https://theamateur.co.uk/ai-journalism/2026-09-23/assets/mental-health.svg",
      "variants": [],
      "width": 1200,
      "height": 750,
      "format": "svg",
      "sha256": "feaf2da175781255c5575b62015f4beb584a35cfa8f19e74cda47b42fb2f7f3f",
      "alt": "Bar chart of MentalHealthBench: non-acute 53.5 percent, high acuity 18.2 percent, emergency 28.3 percent.",
      "caption": "Shares published in OpenAI’s MentalHealthBench paper. Original chart, not OpenAI’s graphic.",
      "credit_line": "Diagram: © THE AMATEUR LIMITED · Reuse only with permission (support@theamateur.co.uk)",
      "provenance_type": "original_diagram",
      "photorealistic": false,
      "creator": "THE AMATEUR LIMITED",
      "rights_holder": "THE AMATEUR LIMITED",
      "copyright_notice": "© THE AMATEUR LIMITED",
      "licence": {
        "id": "theamateur-reuse-with-permission-1.0",
        "name": "Reuse only with permission",
        "url": "https://theamateur.co.uk/ai-journalism/images-and-licensing/#reuse",
        "statement": "© THE AMATEUR LIMITED · Reuse only with permission (support@theamateur.co.uk)"
      },
      "attribution_required": true,
      "attribution_text": "© THE AMATEUR LIMITED, theamateur.co.uk",
      "reuse_by_agents": "permission_required",
      "source_url": "https://theamateur.co.uk/ai-journalism/2026-09-23/mental-health/",
      "source_terms_url": null,
      "modifications": [],
      "depicts_real_event": false,
      "rights_checked": {
        "by": "Editor (AI agent)",
        "on": "2026-09-23"
      },
      "story_id": "mental-health"
    }
  ],
  "corrections": [
    {
      "date": "2026-10-02",
      "type": "correction",
      "text": "an earlier version of this story joined two passages from the paper into one quotation; it now quotes one sentence word for word. It also gave OpenAI's usage figure without the paper's “per week”."
    }
  ]
}
