🧠 OpenAI releases open benchmark for evaluating AI in psychology
The MentalHealthBench dataset is under an MIT license: 1,215 synthetic psychological dialogues and 5,262 expert criteria — adults, teens 13–17, caregivers, and doctors, 11 languages. The criteria were written by more than 80 psychologists and psychiatrists from 22 countries, and the GPT-5.6 Sol grader evaluates them.
🌍 The first open benchmark measures AI behavior in mental health across the full spectrum of acuity, not just emergency scenarios. Open weights and methodology allow models to be re-verified, and weighted criteria from −10 to +10 penalize harmful responses.
👤 The dataset (2.1 MB JSONL with a canary string) and 30-page report are open. The leader GPT-6 Astra has only 57.3% ideal behaviors, GPT-6 Sol — 53.9%, Claude Opus 5.5 — 52.4%: top bots still fall short of clinical guidelines.
Source 1: https://openai.com/index/introducing-mentalhealthbench/ Source 2: https://cdn.openai.com/ctf-cdn/OAI_MentalHealthBench.zip
