Back to Search View Original Cite This Article

Abstract

<p>Large language models (LLMs) are increasingly used for mental health support, yet standards for evaluating their safety and competence remain unsettled. Benchmarks, standardized tests with predetermined scoring criteria, are the most common evaluation approach, but this landscape has not been systematically mapped. In this systematic scoping review, we searched four databases, machine-learning repositories, and grey literature through February 2026 and identified 173 publicly available mental health benchmarks for LLMs, extracting 48 fields covering what each benchmark measures, how it was constructed, and how performance is evaluated. The field has grown sharply recently, with more than 70% of benchmarks appearing in 2024 or later, but remains shaped by data availability rather than clinical relevance. Risk and safety, diagnosis, and therapeutic skills dominate, while bias and stigma are nearly absent. Social media and LLM-simulated content are the leading data sources, and data drawn from clinical care are rare. Ethical and methodological safeguards, including ethics review, demographic reporting, human baselines, and hidden test sets, were reported infrequently. Nearly half of the benchmarks involve an LLM in generating items, annotating labels, or judging outputs, a share that reached 90% of new releases by 2026. Shared reporting standards, clinically grounded data, and stronger safeguards are needed for benchmark scores to support claims about safety and clinical capability in mental health care.</p>

Show More

Keywords

benchmarks data mental health safety

Related Articles

PORE

About

Connect