Joel Scanlon Digital Specialist and Founder of News-Medical.Net | Official Website
+ Pharmaceuticals
Patient Daily | Jul 12, 2026

New benchmark shows AI models fall short in emotional insight for mental health care

A new benchmark reveals that advanced artificial intelligence models, while capable of sounding convincing, still lack the emotional insight and clinical judgment needed for reliable mental health care, according to a study published in npj Mental Health Research on July 12.

Researchers developed PsyEval, a comprehensive evaluation tool designed to assess large language models in three areas: knowledge understanding, diagnosis and assessment, and emotional support. The benchmark was used to evaluate eleven leading language models across these dimensions. Findings indicate that although some models can mirror human fluency and demonstrate strong factual recall, they struggle with deeper emotional understanding and nuanced clinical tasks.

PsyEval results showed that evaluated models exhibited high prompt sensitivity but had an "empathy gap" compared to human counselors. The study noted that safety-related refusals by some large foundation models could lower their diagnostic classification scores, while less guarded models might risk overdiagnosis by classifying patients more aggressively. For example, larger general-purpose language models performed better on factual knowledge tasks; Qwen2.5-72B achieved 91% accuracy on the Chinese MCMLE-mental dataset, while GPT-4-turbo led English tasks with about 76% accuracy. However, performance dropped during urgent psychiatric scenarios—GPT-4-turbo's accuracy fell to around 73% on crisis response questions.

The research also identified an "inverse scaling" phenomenon where smaller models like LLaMa-3-8B outperformed larger ones such as GPT-4 in certain diagnostic classifications—for instance, achieving near-perfect accuracy for attention-deficit/hyperactivity disorder (ADHD) at 100%, compared to approximately 25% for GPT-4. The authors said this may reflect a trade-off between safety guardrails and utility: stricter guardrails can lead to more refusals or ambiguous responses counted as incorrect under PsyEval criteria.

When it came to providing emotional support, humans consistently outperformed AI systems in asking probing questions during counseling interactions; humans achieved higher “Exploration scores” than all evaluated AI systems across both English and Chinese datasets. Prompt style was also found to significantly influence model performance—scenario simulation prompts generally produced higher empathy scores than step-by-step reasoning prompts.

The authors said their study did not directly compare model performance against trained professionals, nor test outcomes with real patients or autonomous crisis intervention scenarios. They concluded that future development should focus on incorporating culturally diverse alignment data and smarter safety mechanisms so that user protection does not come at the expense of potential clinical utility.

Organizations in this story

More News