Ian Birkby CEO | News Medical
+ Pharmaceuticals
Patient Daily | Jun 21, 2026

Google’s AMIE outperforms doctors in simulated multi-visit disease management study

Google’s Articulate Medical Intelligence Explorer (AMIE) matched or outperformed primary care physicians on several key disease-management reasoning tasks in a blinded virtual study, according to research published as an Accelerated Article Preview in the journal Nature on June 21. The system remains experimental and untested in real clinical care, researchers said.

The Google DeepMind and Google Research team evaluated AMIE, a large language model-based artificial intelligence system designed for conversational diagnostic and disease management tasks over multiple patient visits. To improve its performance for management reasoning, the team developed an agentic system that includes an empathetic dialogue agent for synchronous text-chat with patients and a management reasoning agent that references up-to-date clinical guidelines and drug formularies.

Researchers benchmarked medication reasoning using RxQA, a multiple-choice question set derived from OpenFDA and the British National Formulary, validated by board-certified pharmacists. In randomized, blinded comparisons across 100 multi-visit case scenarios based on UK NICE Guidance and BMJ Best Practice guidelines, AMIE was found to be non-inferior to physicians overall. It scored significantly higher than doctors on appropriateness of treatment plans, precision of investigations recommended, avoidance of significant errors at some visits, follow-up recommendations, and guideline alignment.

Medication reasoning accuracy was assessed under both “open-book” (with access to external information) and “closed-book” conditions. Both AMIE and physicians benefited from external resources; however, AMIE outperformed doctors on more difficult questions regardless of access.

Researchers highlighted that while LLM-based systems like AMIE show promise as future tools for managing chronic diseases—potentially providing continuity in fragmented health systems—they cautioned that the technology is not ready for clinical use. The study used simulated consultations with trained actors rather than real patients or settings; scenarios were constructed specifically for evaluation purposes rather than reflecting routine primary care; effects on actual patient outcomes were not measured.

The authors emphasized rapid advancements in LLMs could help address limitations such as confabulations, but urged further comprehensive research before deployment: “the researchers urge that it be seen as a first step in measuring management reasoning and highlight the need for future work to explore the reasoning traces of medical AI systems in a comprehensive, quantitative manner.”

Organizations in this story