16 Sources
[1]
In Harvard study, AI offered more accurate diagnoses than emergency room doctors | TechCrunch
A new study examines how large language models perform in a variety of medical contexts, including real emergency room cases -- where at least one model seemed to be more accurate than human doctors. The study was published this week in Science and comes from a research team led by physicians and
[2]
AI can reason like a physician -- what comes next?
Large language models (LLMs) are artificial intelligence (AI) algorithms that are trained on vast amounts of data to learn patterns that enable them to generate human-like responses. Reasoning models are LLMs with the added capability of working through problems step by step before responding, thus
[3]
AI Outperforms ER Doctors in Diagnostic Cases, Study Points to Collaborative Care
Macy has been working for CNET for coming on 2 years. Prior to CNET, Macy received a North Carolina College Media Association award in sports writing. Have you ever thought about how artificial intelligence compares to a human physician in an emergency diagnostic setting? New research published
[4]
Can AI help doctors avoid missed diagnoses? A new study suggests yes
Humans still have important roles to play in medicine, experts stress In some of medicine's toughest cases, the hardest part isn't choosing the right diagnosis. It's thinking of it at all. Artificial intelligence may now be better at that than doctors, a new study suggests. "We're witnessing a
[5]
AI in the emergency department: promising, powerful but still unproven
Artificial intelligence can now outperform doctors at diagnosing patients in the emergency department, according to a new study in Science. The AI was given written notes from real emergency department records from a hospital in Boston, US, and asked to weigh in at different points during the
[6]
Large language model outperforms human doctors in clinical reasoning tasks
American Association for the Advancement of Science (AAAS)Apr 30 2026 A cutting-edge large language model (LLM) outperformed human doctors in common clinical reasoning tasks including emergency room decisions, identifying likely diagnoses, and choosing next steps in management, according to a new
[7]
AI Just Beat Doctors at Diagnosing ER Patients. Don't Get All Excited
Emergency departments and other clinical settings across the world are now one step closer to sounding like the cockpit of the Millennium Falconâ€"with human doctors soliciting advice from, bickering with, and not infrequently trusting the guidance of their opinionated AI colleagues. Researchers
[8]
In real-world test, an AI model did better than ER doctors at diagnosing patients
Researchers tested an AI model against ER doctors and found the model outperformed the humans. shapecharge/E+/Getty Images hide caption A patient shows up at the hospital with a pulmonary embolism -- a blood clot that has traveled to the lungs. After initially improving, their symptoms start to
[9]
Medical AI matches doctors on diagnoses, but that doesn't make it safe
Advanced medical artificial intelligence (AI) is improving fast. In difficult diagnosis tests built from real patient cases, some AI systems can now match - and sometimes outperform - experienced doctors. That progress is raising larger questions about safety and patient care. A system that
[10]
AI outperforms doctors in Harvard trial of emergency triage diagnoses
Researchers say results mark a 'profound change in technology that will reshape medicine' From George Clooney in ER to Noah Wyle in The Pitt, emergency department doctors have long been popular heroes. But will it soon be time to hang up the scrubs? A groundbreaking Harvard study has found that
[11]
Study: AI can outperform doctors on diagnosing cases
AI performed as well or better than physicians in new study. Credit: Rawlstock via Moment / Getty Images Artificial intelligence that can "reason" is now capable of diagnosing real-life medical scenarios as well as or better than physicians, according to the results of a study published Thursday
[12]
AI nailed emergency diagnoses better than doctors in Harvard trials
AI beat doctors in emergency triage, but don't fire your physician yet AI has plenty of messy use cases, but emergency medicine may be one place where it can do some real good. A Harvard study comparing AI performance against doctors using patient data from emergency-room cases revealed that
[13]
A major new study found AI outperformed doctors in ER diagnosis -- but there's a catch
"No one should look at this and say we do not need doctors," Rodman said in a call with reporters. At the same time, the researchers did argue that AI had reached the point where it could be a genuine asset for doctors in certain situations -- especially in the ER, where physicians are frequently
[14]
AI outperforms doctors in real ER tests, raising safety questions
In a recent study, an AI system beat physicians across a broad set of medical reasoning tests, including messy emergency room cases drawn from real records. The result pushes medical AI beyond exam success and toward the harder question of whether it can be tested safely in hospitals. In 76
[15]
Landmark Test of Clinical Reasoning Finds AI Outperformed | Newswise
AI can pass the hardest exams medical school has to offer. But can it handle the real world's inherent messiness? Harvard Medical School and Beth Israel Deaconess Medical Center researchers sought to find out. Newswise -- BOSTON - In one of the largest studies to compare artificial intelligence
[16]
Harvard Researchers Say Your Next ER Diagnosis May Come From AI -- and It Could Be More Accurate Than Human Doctors
What if the most accurate diagnosis in the emergency room didn't come from a doctor -- but AI? A new study led by a team of researchers at Harvard Medical School suggests that moment may already be here. The study, which was published in Science, a peer-reviewed journal, found that AI systems
Share
Copy Link
A Harvard Medical School study published in Science reveals that OpenAI's o1 model achieved 67% diagnostic accuracy at emergency room triage, outperforming two attending physicians who scored 55% and 50%. The research tested AI against doctors using real patient data from Beth Israel Deaconess Medical Center, with blinded reviewers unable to distinguish AI output from human diagnoses. While the findings demonstrate AI's potential as a diagnostic tool, researchers emphasize the urgent need for clinical trials and accountability frameworks before deployment in actual patient care settings.
A groundbreaking study from Harvard Medical School and Beth Israel Deaconess Medical Center demonstrates that AI diagnosis capabilities now match or exceed physician performance in emergency medicine settings. Published in Science, the research tested OpenAI's o1 and 4o models against human doctors using 76 actual emergency room cases, marking a shift from theoretical assessments to authentic clinical evaluation
1
.
Source: Inc.
The results show the o1 model achieved the exact or very close diagnosis in 67% of triage cases, while two attending physicians scored 55% and 50% respectively
1
. This performance gap proved most pronounced at initial triage, the critical moment when information is scarcest and decisions carry the highest urgency. Blinded reviewers assessing the diagnoses could not distinguish between AI-generated and human recommendations3
.The research team conducted six experiments to measure clinical diagnostic reasoning across multiple scenarios. When tested on published clinicopathological conference cases, the o1-preview model achieved exact or very-close diagnostic accuracy in 88.6% of cases, substantially outperforming GPT-4 which scored 72.9%
2
. This advancement in large language model diagnostic accuracy represents a significant leap from earlier AI systems that primarily demonstrated proficiency on medical licensing examinations rather than real-world patient care.Arjun Manrai, who heads an AI lab at Harvard Medical School and serves as one of the study's lead authors, stated: "We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines"
1
. The researchers emphasized they did not pre-process patient data, presenting the AI with the same information available in electronic medical records at each diagnostic touchpoint1
.The study found AI as a diagnostic tool handled uncertainty far more effectively than human clinicians, particularly when working with fragmented or unstructured health data and notes
3
. Adam Rodman, a Beth Israel doctor and lead author, described a case where a patient with routine respiratory symptoms who had recently undergone organ transplant turned out to have a dangerous flesh-eating infection. "The model actually was suspicious of this [infection] from the very beginning, probably 12 to 24 hours before the human physician would have become suspicious"4
.Both human clinicians and AI improved as more patient data became available, but the model's advantage at early stages suggests potential for avoiding missed diagnoses with AI support
3
. This capability addresses one of emergency medicine's most challenging aspects: thinking of the correct diagnosis when information is limited and time is critical4
.
Source: Earth.com
Despite the impressive results showing AI outperforms ER doctors in specific contexts, researchers stress the technology should augment rather than replace physicians. "I don't think our findings mean that AI replaces doctors, despite what some companies are likely to say, and how they're likely to use these results," Manrai said during a press briefing
3
. Rodman told the Guardian that patients "want humans to guide them through life or death decisions [and] to guide them through challenging treatment decisions"1
.Previous research on collaborative care models found no substantial difference between physicians augmented with GPT-4 and the GPT-4 model working alone, though both outperformed physicians with conventional resources
2
. This suggests determining optimal implementation will require evaluating AI alone, clinician alone, and clinician with AI configurations2
.Related Stories
The study identifies an urgent need for prospective clinical trials to evaluate AI in healthcare within real-world patient care settings
1
. Currently, "there's no formal framework right now for accountability" around AI diagnoses, according to Rodman1
. Researchers at Flinders University wrote in a Science commentary that "we do not allow doctors to practice without supervision and evaluation, and AI should be held to comparable standards"3
.The research carries important limitations. The models only processed text-based information, while actual emergency medicine relies heavily on visual and auditory cues from physical examinations
1
. The AI never saw patients, examined them, spoke to families, or took responsibility for outcomes5
. Future assessments must evaluate multimodal AI capabilities that process images, audio, and video alongside text2
.
Source: CNET
The urgency for establishing safety and equity standards intensifies as AI adoption accelerates. A Royal College of Physicians survey found 16% of UK doctors use AI tools in clinical practice daily, with another 15% using them weekly
5
. Globally, 1 in 5 doctors and nurses used AI for second opinions on complex cases as of 2025, with over half wanting to use it for this purpose4
.Doctors are integrating these tools into practice, sometimes without institutional oversight, before hospitals have established protocols for assessment, staff training, harm detection, or decision support accountability
2
. The gap between producing possible diagnoses and actually improving patient outcomes remains unclear, as longer diagnostic lists could generate unnecessary tests, over-treatment, or unwarranted confidence in plausible but incorrect answers5
. Regulators, hospitals, and healthcare providers must collaborate to test these tools thoroughly, ensuring they deliver care that is better, safer, and faster for all patients3
.Summarized by
Navi
[1]
[2]
[5]
01 Oct 2025•Health

05 Apr 2025•Health

18 Aug 2025•Health

1
Technology

2
Science and Research

3
Technology
