3 Sources
[1]
How good are AI doctors at medical conversations?
Artificial intelligence tools such as ChatGPT have been touted for their promise to alleviate clinician workload by triaging patients, taking medical histories and even providing preliminary diagnoses. These tools, known as large-language models, are already being used by patients to make sense of
[2]
AI models struggle in real-world medical conversations
Harvard Medical SchoolJan 2 2025 Artificial intelligence tools such as ChatGPT have been touted for their promise to alleviate clinician workload by triaging patients, taking medical histories and even providing preliminary diagnoses. These tools, known as large-language models, are already being
[3]
AI chatbots fail to diagnose patients by talking with them
Although popular AI models score highly on medical exams, their accuracy drops significantly when making a diagnosis based on a conversation with a simulated patient Advanced artificial intelligence models score well on professional medical exams but still flunk one of the most crucial physician
Share
Copy Link
A new study reveals that while AI models perform well on standardized medical tests, they face significant challenges in simulating real-world doctor-patient conversations, raising concerns about their readiness for clinical deployment.

A groundbreaking study led by researchers from Harvard Medical School and Stanford University has revealed a significant gap between the performance of AI models in standardized medical tests and their ability to handle real-world patient interactions. The research, published in Nature Medicine, introduces a new evaluation framework called CRAFT-MD (Conversational Reasoning Assessment Framework for Testing in Medicine) designed to assess the capabilities of large language models in medical settings
1
.While AI tools like ChatGPT have shown promise in alleviating clinician workload through patient triage and preliminary diagnoses, the study exposes a striking paradox. Dr. Pranav Rajpurkar, assistant professor of biomedical informatics at Harvard Medical School, notes, "While these AI models excel at medical board exams, they struggle with the basic back-and-forth of a doctor's visit"
2
.The CRAFT-MD framework simulates real-world interactions by evaluating how well large language models can collect patient information and make diagnoses. It employs AI agents to pose as patients and grade the accuracy of diagnoses, with human experts providing additional evaluation
1
.The study tested four AI models, including both proprietary and open-source versions, across 2,000 clinical vignettes. The results showed a significant drop in performance when models engaged in conversational, open-ended interactions compared to answering multiple-choice questions
3
.Key findings include:
1
.Related Stories
The research team offers several recommendations for AI developers and regulators:
2
.While the study highlights current limitations, it also paves the way for more robust AI tools in healthcare. Dr. Rajpurkar emphasizes that even if AI models improve, they would likely serve as powerful support tools rather than replacements for experienced physicians
3
.Summarized by
Navi
[1]
[2]
[3]
05 Mar 2026•Health

30 Apr 2026•Science and Research

09 Feb 2026•Health

1
Technology

2
Technology

3
Technology
