9 Sources
[1]
ChatGPT Health fails critical emergency and suicide safety tests
Mount Sinai Health SystemFeb 24 2026 ChatGPT Health, a widely used consumer artificial intelligence (AI) tool that provides health guidance directly to the public-including advice about how urgently to seek medical care-may fail to direct users appropriately to emergency care in a significant
[2]
ChatGPT Health misses urgent medical crises over 50% of the time
While OpenAI claims continuous model refinement and disputes the study's real-world applicability, the research highlights current limitations in AI medical assessment tools. According to new research published in Nature Medicine, ChatGPT Health (OpenAI's dedicated AI-driven chatbot that's
[3]
'Unbelievably dangerous': experts sound alarm after ChatGPT Health fails to recognise medical emergencies
ChatGPT Health regularly misses the need for medical urgent care and frequently fails to detect suicidal ideation, a study of the AI platform has found, which experts worry could "feasibly lead to unnecessary harm and death". OpenAI launched the "Health" feature of ChatGPTto limited audiences in
[4]
ChatGPT Health 'under-triaged' half of medical emergencies in a new study
ChatGPT Health -- OpenAI's new health-focused chatbot -- frequently underestimated the severity of medical emergencies, according to a study published last week in the journal Nature Medicine. In the study, researchers tested ChatGPT Health's ability to triage, or assess the severity of, medical
[5]
Research Identifies Blind Spots in AI Medical Triage | Newswise
Newswise -- New York, NY [February 24, 2026] -- ChatGPT Health, a widely used consumer artificial intelligence (AI) tool that provides health guidance directly to the public -- including advice about how urgently to seek medical care -- may fail to direct users appropriately to emergency care in a
[6]
What to know before asking an AI chatbot for health advice
WASHINGTON (AP) -- With hundreds of millions of people turning to chatbots for advice, it was only a matter of time before tech companies began offering programs specifically designed to answer health questions. In January, OpenAI introduced ChatGPT Health, a new version of its chatbot that the
[7]
ChatGPT Health fails to spot 52% of medical emergencies in study
A study published in Nature Medicine on February 24 found that ChatGPT Health failed to direct users to emergency care in more than half of serious medical cases. Researchers at the Icahn School of Medicine at Mount Sinai conducted the evaluation, testing the consumer-facing tool across 960
[8]
What to Know Before Asking an AI Chatbot for Health Advice
This image provided by OpenAI in February 2026 demonstrates a health chatbot on a phone app. (OpenAI via AP) WASHINGTON (AP) -- With hundreds of millions of people turning to chatbots for advice, it was only a matter of time before tech companies began offering programs specifically designed to
[9]
Where ChatGPT Health fails -- and how it could turn deadly
OpenAI last month introduced ChatGPT Health, a dedicated space in ChatGPT that allows users to ask health questions, analyze their medical records and connect to wellness apps. Now, weeks after its launch, researchers from the Icahn School of Medicine at Mount Sinai are raising concerns that the
Share
Copy Link
OpenAI's ChatGPT Health missed over half of medical emergencies in a Nature Medicine study, directing patients to routine appointments instead of emergency rooms. With 40 million daily users seeking health guidance, the AI tool also showed alarming inconsistencies in suicide-crisis safeguards, triggering alerts for low-risk cases while failing to respond when users described specific self-harm plans.

ChatGPT Health, OpenAI's dedicated consumer AI tool for health guidance, failed to appropriately direct users to emergency care in more than half of serious medical emergencies, according to the first independent safety evaluation published in Nature Medicine
1
. The study, conducted by researchers at Mount Sinai Health System, tested 60 realistic patient scenarios spanning 21 medical specialties across 960 interactions and found that the AI for health guidance under-triaged 51.6% of cases that physicians determined required immediate emergency care3
. Instead of recommending emergency room visits, OpenAI's ChatGPT Health directed patients experiencing life-threatening conditions like diabetic ketoacidosis and respiratory failure to schedule routine appointments within 24 to 48 hours4
.The stakes are extraordinarily high given that approximately 40 million people use the tool daily to seek health information and decide whether to seek urgent medical crises care
1
. Lead author Ashwin Ramaswamy, an Instructor of Urology at the Icahn School of Medicine, explained that the study aimed to answer a basic but critical question: "If someone is experiencing a real medical emergency and turns to ChatGPT Health for help, will it clearly tell them to go to the emergency room?"5
.While ChatGPT Health performed well in textbook emergencies such as stroke or severe allergic reactions—correctly triaging these 100% of the time—it struggled significantly in more nuanced situations where the danger is not immediately obvious
4
. In one asthma scenario, the system identified early warning signs of respiratory failure in its own explanation but still advised waiting rather than seeking emergency treatment1
. This paradoxical behavior reveals critical AI blind spots in clinical judgment where the language model (LLM) recognizes dangerous findings yet still reassures patients.Doctoral researcher Alex Ruani, who studies health misinformation mitigation at University College London, described the findings as "unbelievably dangerous," noting that in one simulation, the platform sent a suffocating woman to a future appointment she wouldn't live to see eight times out of 10 attempts—an 84% failure rate
3
. Meanwhile, the chatbot guidance also over-triaged 64.8% of nonurgent cases, recommending immediate medical care for completely safe individuals2
.The independent safety evaluation revealed particularly concerning failures in suicide-crisis safeguards designed to direct users to the 988 Suicide and Crisis Lifeline in high-risk situations
1
. Researchers found that these alerts appeared inconsistently and were "inverted relative to clinical risk," appearing more reliably for lower-risk scenarios while failing to appear when users described specific plans for self-harm5
.In testing suicidal ideation scenarios, Ramaswamy described a case where a 27-year-old patient described thinking about taking a lot of pills. When the patient described symptoms alone, the crisis intervention banner linking to suicide help services appeared every time. However, when normal lab results were added to the same patient scenario with identical words and severity, the banner vanished—zero out of 16 attempts
3
. Senior study author Girish N. Nadkarni, Chief AI Officer of the Mount Sinai Health System, noted that "when someone talks about exactly how they would harm themselves, that's a sign of more immediate and serious danger, not less"1
.The research team created 60 structured clinical scenarios covering conditions from mild illnesses to true medical emergencies. Three independent physicians determined the correct level of urgency for each case using guidelines from 56 medical societies
5
. Each scenario was tested under 16 different contextual conditions, including variations in race, gender, social dynamics such as someone minimizing symptoms, and barriers to care like lack of insurance or transportation1
. The variations were designed to produce the exact same triage recommendations regardless of demographic changes, and the study found no significant differences based on these factors4
.Notably, the platform was nearly 12 times more likely to downplay symptoms when the "patient" mentioned a "friend" in the scenario suggested it was nothing serious
3
. This susceptibility to social influence demonstrates how the consumer AI tool fails to recognize medical emergencies when contextual noise is introduced.Related Stories
Isaac S. Kohane, Chair of the Department of Biomedical Informatics at Harvard Medical School, who was not involved with the research, emphasized that "LLMs have become patients' first stop for medical advice—but in 2026 they are least safe at the clinical extremes, where judgment separates missed emergencies from needless alarm. When millions of people are using an AI system to decide whether they need emergency care, the stakes are extraordinarily high. Independent evaluation should be routine, not optional"
1
.Prof. Paul Henman, a digital sociologist at the University of Queensland, warned that if ChatGPT Health was used by people at home, "it could lead to higher numbers of unnecessary medical presentations for low-level conditions, and a failure of people to obtain urgent medical care when required, which could feasibly lead to unnecessary harm and death"
3
. He also raised concerns about legal liability, noting that a suite of legal cases against tech companies are already in motion related to suicide and self-harm after using AI chatbots.An OpenAI spokesperson told media outlets that while the company welcomed independent research evaluating AI systems in healthcare, the study did not reflect how people typically use ChatGPT Health in real life
3
. The spokesperson emphasized that the model is continuously updated and refined, and that the chatbot is designed for users to ask follow-up questions to provide more context rather than give single responses to medical scenarios4
. ChatGPT Health is currently available only to a limited number of users on a waitlist, and OpenAI is working to improve safety and reliability before wider release.However, researchers and experts maintain that a plausible risk of harm is sufficient to justify stronger safeguards and independent oversight
3
. Dr. John Mafi, an associate professor of medicine at UCLA Health, stressed that "before you roll something like this out, to make life-affecting decisions, you need to rigorously test it in a controlled trial, where you're making sure that the benefits outweigh the harms"4
. The study authors advise that for worsening or concerning symptoms, including chest pain, shortness of breath, severe allergic reactions, or changes in mental status, people should seek medical care directly rather than relying solely on AI tools.Summarized by
Navi
[1]
[3]
05 Mar 2026•Health

14 Apr 2026•Health

24 Jul 2026•Technology

1
Technology

2
Science and Research

3
Technology
