17 Sources
[1]
Americans ask AI for health care. Hospitals think the answer is more chatbots.
With many Americans turning to Large Language Models for health advice, health systems around the country are eyeing and even rolling out their own branded chatbots in an attempt to harness this already popular tool and steer more people to their services. But the burgeoning trend is raising
[2]
AI Chatbots Give Misleading Medical Advice 50% of the Time, Study Finds
Artificial intelligence-driven chatbots are giving users problematic medical advice about half the time, according to a new study, highlighting the health risks of the technology that's becoming increasingly integral in day-to-day life. Researchers from the US, Canada and the UK evaluated five
[3]
AI Chatbots Vary Widely on Fake Medical References
Hallucination rates for medical references generated by artificial intelligence (AI) chatbots range from zero to more than a third, depending on the platform, a new study has found -- with some tools performing reliably and others producing large numbers of fabricated citations. The study,
[4]
AI chatbots misdiagnose in over 80% of early medical cases, study finds
Consumer AI chatbots falter when used to make medical diagnoses, particularly when faced with incomplete information, according to new research highlighting the risks of relying on them as digital doctors. The study finds that leading large language models struggle to suggest a range of possible
[5]
Study finds popular AI chatbots often give problematic health advice
By Hugo Francisco de SouzaReviewed by Susha Cheriyedath, M.Sc.Apr 16 2026 A new audit suggests that widely used free AI chatbots can sound confident while delivering misleading health information, weak citations, and advice that may be unsafe without expert guidance. Study: Generative artificial
[6]
Why many Americans are turning to AI for health advice, according to recent polls
NEW YORK (AP) -- When Tiffany Davis has a question about a symptom from the weight-loss injections she's taking, she doesn't call her doctor. She pulls out her phone and consults ChatGPT. "I'll just basically let ChatGPT know my status, how I'm feeling," said the 42-year-old in Mesquite, Texas. "I
[7]
Can I trust health advice from an AI chatbot?
For the past year, Abi has been using ChatGPT - one of the best known AI chatbots - to help manage her health. The appeal is clear. It can feel impossible to get hold of a GP and artificial intelligence is always ready to answer your questions. And AI has comfortably passed some medical exams. So
[8]
Ready or Not, LLMs Are Coming for Medicine
This transcript has been edited for clarity. Welcome to Impact Factor, your weekly dose of commentary on a new medical study. I'm Dr F. Perry Wilson from the Yale School of Medicine. There's a new genre of medical papers in the "AI in medicine" space, and, like Mulder from The X-Files, I want to
[9]
Generative AI falls short in diagnostic reasoning despite accuracy
Mass General BrighamApr 13 2026 Despite increasing use of artificial intelligence (AI) in health care, a new study led by Mass General Brigham researchers from the MESH Incubator shows that generative AI models continue to fall short at their clinical reasoning capabilities. By asking 21
[10]
Millions of Americans Are Talking to AI Instead of Going to the Doctor, and It's Giving Them Horrendously Flawed Medical Advice
Can't-miss innovations from the bleeding edge of science and tech While Google's AI may no longer recommend eating rocks or confidently telling users to put glue on their pizza, even cutting-edge AI chatbots remain staggeringly incompetent at dispensing medical advice. In a new study published
[11]
Millions of Americans are talking to AI about health, and some are dangerously skipping real doctors
One in four Americans already relies on AI for health advice, a trend that raises serious concerns. Google used to be the go-to service for people who wanted to learn about their health conditions. The tide has been slowly shifting with more and more users turning to AI for their health-related
[12]
Americans Turning to AI to Supplement Healthcare Visits
Editor's Note: This research was conducted in partnership with West Health through the West Health-Gallup Center on Healthcare in America, a joint initiative to report the voices and experiences of Americans within the healthcare system. WASHINGTON, D.C. -- As artificial intelligence becomes
[13]
AI fails at primary diagnosis more than 80% of the time, study finds
Generative artificial intelligence (AI) still lacks the reasoning processes needed for safe clinical use, a new study has found. AI chatbots have improved their diagnostic accuracy when presented with comprehensive clinical information, but still failed to produce an appropriate differential
[14]
New findings from this Gallup poll show how Americans are using AI for health advice
Most recent AI health users are looking for quick answers Most Americans using AI tools for health purposes say they want immediate answers. In some cases, it helps them evaluate what kind of medical attention they need. "It'll let me know if something's serious or not," Davis said of ChatGPT,
[15]
ChatGPT And Copilot Are Becoming Americans' First Stop For Medical Questions -- But Trust Still Lags - Alph
A new national survey released Tuesday shows that 25% of American adults have used an artificial intelligence (AI) tool or chatbot to seek health information or advice, underscoring how AI is becoming embedded in healthcare decision-making as a supplemental tool rather than a replacement for
[16]
AI Chatbots Can Diagnose. Doctors Have Questions. | PYMNTS.com
By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions. One of the most visible expressions of that promise is the proliferation of AI
[17]
AI doc bots fall for fake disease -- and diagnose folks with it
Swedish researchers fed a fake medical diagnosis, along with phony scientific studies, into AI chatbots to see if they would fall for it - and they did. A team led by Almira Osmanovic Thunström at the University of Gothenburg cooked up a completely fraudulent eye condition called bixonimania -- a
Share
Copy Link
Multiple studies reveal AI chatbots deliver problematic health advice 50% of the time, with some platforms fabricating medical references at rates up to 34%. As one in three Americans now turn to AI for health information, researchers warn these tools lack clinical judgment and produce authoritative-sounding but potentially dangerous responses, especially when patient data is incomplete.

As one in three Americans turn to AI chatbots for health information, multiple studies reveal a troubling pattern: these AI tools for health advice deliver misleading medical advice at rates that should concern anyone using them. A study published in BMJ Open found that five popular platforms—ChatGPT, Gemini, Meta AI, Grok, and DeepSeek—produced problematic health advice in approximately 50% of cases, with nearly 20% deemed highly problematic
2
. The evaluation involved 250 prompts across five misinformation-prone categories including cancer, vaccines, stem cells, nutrition, and athletic performance5
.The consumer AI chatbots performed relatively better on closed-ended prompts and questions related to vaccines and cancer, but struggled significantly with open-ended prompts and domains like nutrition. Open-ended questions generated 40 highly problematic responses compared to just 9 for closed-ended prompts
5
. Critically, these Large Language Models (LLMs) delivered answers with confidence and certainty despite their flaws, creating a dangerous illusion of reliability that could compromise patient safety.Beyond inaccurate advice, AI chatbots fabricate medical references at concerning rates. Research published in The Annals of the Royal College of Surgeons of England examined nine AI platforms and discovered hallucination rates ranging from zero to 34% for AI-generated references
3
. Grok 3 performed worst with 34% of references fabricated or unverifiable, while DeepSeek DeepThink followed at 25%. Only five of the nine models tested produced no hallucinated references at all.The most concerning fake medical references "closely resembled legitimate scientific literature," featuring plausible article titles, invented URLs, and attributions to reputable institutions like the Mayo Clinic
3
. This sophisticated fabrication undermines users' ability to verify whether information is accurate or evidence-based. No chatbot in the BMJ Open study produced a fully complete and accurate reference list in response to any prompt2
. Additionally, many cited sources were behind academic paywalls, further limiting verification capabilities—though Google Gemini stood out by providing all open-access, directly clickable sources3
.When it comes to diagnosis, AI for health care faces even steeper challenges. A study published in Jama Network Open tested 21 LLMs using clinical vignettes and found that failure rates exceeded 80% for all models when performing differential diagnosis with incomplete patient information
4
. The models from OpenAI, Anthropic, Google, xAI, and DeepSeek struggled particularly at the open-ended start of cases when limited data was available. "These models are great at naming a final diagnosis once the data is complete, but they struggle at the open-ended start of a case, when there isn't much information," said lead author Arya Rao4
.Failure rates dropped below 40% for final diagnoses with complete data, with top performers exceeding 90% accuracy
4
. However, this highlights a critical limitation: real-world users often input vague or patchy information, precisely the scenario where these tools fail most dramatically. A February study in Nature Medicine involving nearly 1,300 participants found that when researchers provided specific medical scenarios, LLMs correctly identified conditions 95% of the time. But when participants used their own prompts for the same scenarios, accuracy plummeted to just one-third of cases1
. "People don't know what they are supposed to be telling the model," explained lead author Andrew Bean from Oxford University1
.Related Stories
Despite mounting evidence of medical misinformation risks, health systems are rolling out their own branded AI chatbots. K Health is partnering with Hartford HealthCare in Connecticut to deploy its PatientGPT chatbot to tens of thousands of existing patients
1
. CEO Allon Bloch frames this as meeting patients where they are: "Demand is accelerating, and patients are already using AI to navigate their lives"1
.Yet experts question whether sufficient evidence supports these deployments. Adam Rodman, a clinical reasoning researcher at Beth Israel Deaconess Medical Center, told Stat News there isn't yet an evidence base showing that integrating chatbots into health systems improves patient outcomes. "We're not there yet," he said
1
. Concerns extend to monitoring adequacy, liability frameworks, and whether chatbots address the actual care problems patients face. A KFF poll found that among Americans using AI for health queries, 19% cited inability to afford care and 18% lacked a regular provider or couldn't get appointments1
.The explosion of AI chatbots in healthcare occurs against a backdrop of systemic failure. Nearly one-third of Americans—more than 100 million people—lack a primary care provider
1
. OpenAI reports that more than 200 million people ask ChatGPT health and wellness questions weekly2
, while the KFF poll revealed 41% of AI users uploaded personal medical information like test results1
.These tools lack clinical judgment essential for safe medical guidance. Because LLMs generate responses by predicting language patterns rather than retrieving verified facts, they have no built-in mechanism for factual verification
3
. Tim Mitchell, president of the Royal College of Surgeons of England, emphasized that "the excitement around using AI-generated information must be matched with caution, by both patients and doctors"3
. The BMJ Open study authors warned that without public education and oversight, chatbots risk amplifying misinformation through "authoritative-sounding but potentially flawed responses"2
.Watch for regulatory responses addressing liability and transparency requirements. User caution remains essential: verify AI advice with licensed professionals, recognize that premium models may outperform free versions, and understand that confident-sounding prompts don't guarantee accuracy. As specialized medical LLMs like Google's AMIE emerge, their real-world testing with actual patients—particularly in settings with limited doctor access—will determine whether AI can safely supplement rather than substitute for human clinical expertise.
Summarized by
Navi
05 Mar 2026•Health

09 Feb 2026•Health

17 Nov 2025•Health

1
Policy and Regulation

2
Technology

3
Technology
