Ian Birkby CEO | News Medical
+ Pharmaceuticals
Patient Daily | Jun 11, 2026

Article compares AI errors to human confabulations and hallucinations in psychiatry

A recent Perspective article published in the journal NPP–Digital Psychiatry and Neuroscience compared errors made by artificial intelligence systems with confabulations and hallucinations observed in psychiatry, according to a June 11 report. The article discusses how large language models, such as ChatGPT, sometimes generate text that is factually incorrect but appears plausible. Similarly, automatic speech recognition tools like Whisper can produce severe transcription errors under certain conditions, resulting in output that is unfaithful or nonsensical to the source input.

The authors argue that while these AI system errors are often called hallucinations, clarification is necessary. In humans, hallucinations are false sensory experiences occurring without external stimuli. In contrast, large language model errors do not involve perception and are better described as confabulations—false memories produced to fill memory gaps without intent to deceive.

Confabulations typically occur in conditions associated with memory impairment such as Korsakoff’s syndrome and dementia. They range from minor inaccuracies to highly detailed fabrications. The article states that large language models exhibit similar flaws when information is missing due to limitations related to training data or context constraints.

Automatic speech recognition system errors may be more akin functionally—not experientially—to human hallucinations. For example, more than one-third of Whisper's so-called hallucinations contain harmful content including demographic stereotypes or threats of violence. Human auditory verbal hallucinations also often deliver threats or verbal abuse; both systems show repetition of themes or phrases when perceptual signals are weak or degraded.

The article suggests integrating mechanisms into AI systems similar to those used in cognitive-behavioral therapy for humans—such as plausibility assessment using uncertainty estimation methods—to help reduce error rates. Additional strategies include multi-pass verification, retrieval-augmented generation, cross-model verification, semantic entropy methods, prompt design adjustments, and temperature tuning.

The authors conclude that confabulation-like and hallucination-like errors are recognized problems in large language models and related software products. These issues highlight the importance of continued human oversight when using predictive artificial intelligence systems.

Organizations in this story

More News