返回新闻
创新AI Understanding 简报

OpenAI 发布 MentalHealthBench,用于评估心理健康对话中的人工智能

OpenAI 推出了 MentalHealthBench,这是一款由 80 多名临床医生开发的开源评估工具,用于评估 AI 模型如何处理心理健康和情感支持互动。

4 min readRead the linked source
Source-provided image accompanying OpenAI releases MentalHealthBench for evaluating AI in mental health conversations
来源参考来源记录
出版商
edtechinnovationhub.com
来源链接
edtechinnovationhub.comhttps://www.edtechinnovationhub.com/news/openai-releases-mentalhealthbench-to-test-ai-responses-in-mental-health-conversations
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

从这里开始

关键术语

综合数据
用于增强、模拟或保护敏感训练数据的人工生成的数据。
校准
模型的置信度得分与实际正确性概率的匹配程度。
人工智能安全
该领域专注于减少人工智能系统中的有害行为、故障和误用风险。
测试一下自己人工智能道德测验

发生了什么

OpenAI has released MentalHealthBench, an open-source benchmark designed to evaluate how AI models perform during mental health and emotional support conversations. The tool was developed in collaboration with more than 80 licensed psychologists and psychiatrists across 22 countries. It utilizes a dataset of synthetic conversations—covering non-acute stress, high-acuity situations, and emergencies—to test model behaviors such as clinical accuracy, empathy, and urgency . OpenAI reports that it uses its GPT-5.6 Sol model as an automated grader to score responses against expert-defined rubrics.

OpenAI's MentalHealthBench consists of synthetic conversations designed to simulate real-world patterns of AI use in mental health. The dataset includes scenarios involving adults, teenagers, caregivers, and clinicians, spanning 19 languages. The distribution of the benchmark is 53.5% non-acute, 18.2% high-acuity, and 28.3% emergency situations.

The evaluation methodology relies on rubrics created by mental health experts, where responses are scored from -10 to +10 based on clinical importance. Each conversation was reviewed by at least three experts, with criteria retained only upon consensus. OpenAI uses its GPT-5.6 Sol model to automate the grading of other models against these rubrics.

In initial testing, OpenAI reported that GPT-6 Astra achieved the highest score at 57.3%, followed by GPT-6 Sol (53.9%), Claude Opus 5.5 (52.4%), GPT-6 Luna (50.2%), and Muse Spark 1.3 (47%). Older models, including GPT-4o (32.1%) and Gemini 2.5 Pro (29.5%), scored significantly lower.

OpenAI conducted a supplemental study with 44 adults who had prior experience using AI for mental health. The study revealed a divergence in priorities: users prioritized tone and actionable steps, while clinicians prioritized context gathering and the interpretation of ambiguous information. OpenAI noted that these user preferences did not alter the benchmark's expert-derived scoring criteria.

来源详情: edtechinnovationhub.com ↗

为什么这很重要

As AI increasingly becomes a tool for emotional support, establishing standardized, clinically-informed evaluation methods is critical for safety. MentalHealthBench moves beyond simple harm avoidance by measuring nuanced clinical behaviors like context gathering and agency preservation. By providing an open framework, OpenAI allows researchers to compare model performance across a spectrum of mental health needs, highlighting significant performance gaps between newer models like GPT-6 Astra and older iterations like GPT-4o. This transparency is essential for understanding how AI systems navigate the complex, high-stakes domain of mental health, though OpenAI emphasizes that these scores are not measures of clinical effectiveness and that AI remains a supplement, not a replacement, for professional care.

The release of MentalHealthBench addresses a significant gap in : the lack of standardized, clinically-grounded evaluation for emotional support interactions. By moving the conversation beyond binary 'harm avoidance,' the benchmark forces a focus on the quality and accuracy of AI-provided guidance.

The project highlights the tension between clinical standards and user experience. The finding that users and clinicians value different aspects of a conversation underscores the difficulty of designing AI that is both safe and perceived as helpful by individuals in distress.

The benchmark provides a concrete, measurable way to track progress in AI's ability to handle sensitive, multi-turn conversations. As models continue to evolve, this tool serves as a baseline for assessing whether improvements in general reasoning translate into safer, more effective interactions in the mental health domain.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Ethics Quiz

Why can ethical evaluation not be reduced to one model score?

接下来看什么

Researchers and developers can now access the MentalHealthBench dataset to conduct their own evaluations. A key area to monitor is how the industry adopts these specific metrics—such as 'urgency ' and 'clinical accuracy'—in future safety standards. Additionally, the discrepancy between clinician-weighted criteria and user-preferred traits (such as tone and practical steps) suggests that future iterations of benchmarks may need to reconcile expert clinical standards with real-world user expectations. It remains to be seen how other AI providers will respond to these benchmarks and whether they will adopt similar collaborative, expert-led evaluation frameworks.

Watch for whether other major AI labs adopt MentalHealthBench as a standard for their own safety reporting or if they develop competing benchmarks that prioritize different clinical or user-centric metrics.

Monitor the potential for 'benchmark gaming,' where models are specifically optimized to score well on these rubrics without necessarily improving the underlying quality of care or safety in real-world, unscripted interactions.

Observe how the American Psychological Association and other professional bodies respond to the use of and automated grading (via GPT-5.6 Sol) in evaluating clinical-adjacent AI performance.

相关指南和测验

AI 伦理人工智能模型解释AI 的未来人工智能培训测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?