Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Technology

AI Beat Two Doctors on 76 ER Cases—but It Was a Text-Only Test

|Updated: |Author: QUASA Editorial Team|5 min read| 1027
AI Beat Two Doctors on 76 ER Cases—but It Was a Text-Only Test

OpenAI’s o1 produced more accurate diagnoses than two internal medicine attending physicians in a study of 76 emergency-department cases published on April 30, 2026. As of August 2026, that finding remains a retrospective, text-only comparison—not evidence that the model can independently diagnose or treat patients at the bedside.

The result is still important: o1 handled raw clinical notes from real hospital encounters and performed especially well when information was sparse. But subsequent research from members of the same Harvard-led group has sharpened the central warning: models that reason impressively from written records can remain vulnerable when they must reconcile text with images or when the text itself is misleading.

What the 76-case comparison actually measured

The researchers selected records from 76 patients who visited the emergency department at Beth Israel Deaconess Medical Center and were subsequently admitted. OpenAI’s o1 and GPT-4o received the information available in the electronic record at three points: initial triage, the first physician encounter and admission to a hospital floor or intensive care unit.

Two internal medicine attending physicians independently reviewed the same material. Two additional physicians, blinded to whether each response came from a person or a model, then assessed the proposed diagnoses.

At initial triage, o1 supplied the exact diagnosis or a very close alternative in 67% of cases. The two physician comparators reached that standard in 55% and 50% of cases, respectively, according to the detailed 76-case results reported by TechCrunch. The advantage was largest at this earliest stage, when the record contained the least information.

Those numbers describe the quality of a written differential diagnosis under the study’s scoring system. They do not measure whether the model recognized every immediate threat, ordered the right test, chose a safe disposition or improved a patient’s outcome—all central parts of emergency care.

Why “AI beat emergency doctors” is the wrong shorthand

The human comparison involved two internal medicine attendings, not a representative sample of emergency physicians. That distinction matters because emergency medicine is not simply a contest to name the eventual inpatient diagnosis from a chart.

An emergency clinician must stabilize the patient, exclude time-critical conditions, interpret changing symptoms and decide whether discharge, observation or admission is safe. The study’s model did none of those things. It produced an opinion from recorded text after the encounters had occurred.

The cases were real, and the records were not rewritten into polished textbook vignettes. That gives the experiment more practical relevance than a medical multiple-choice examination. Even so, the comparison was retrospective: o1 did not interview patients, perform physical examinations, notice changes in appearance or behavior, communicate uncertainty to families, or bear responsibility for a decision.

The narrow conclusion is therefore stronger than a vague claim that the model merely passed a medical quiz, but much weaker than saying it practiced emergency medicine. It showed unusually capable diagnostic reasoning over messy clinical text.

The later evidence exposes a multimodal weakness

A June 2026 study involving several of the same Harvard researchers tested eight foundation models on 1,090 medical cases combining text and images. Its findings supply a concrete update to the April paper because they examine precisely the kind of nontextual evidence missing from the emergency-record comparison.

The Nature Communications study found that model performance was driven heavily by informative text. In an adversarial experiment, o3 correctly classified 84% of a selected set of images before a misleading fictional vignette was added; accuracy then fell to 28%.

This does not invalidate o1’s result: the studies used different models, datasets and tasks. It does show why strong performance on written records cannot automatically be extended to a patient encounter containing scans, physical signs, physiological signals and possibly inaccurate notes.

The later study also found that adding images to already informative text sometimes failed to improve performance and could reduce it. For clinical systems, the key question is consequently not whether a model can accept several data types, but whether it detects contradictions between them instead of forcing the evidence to fit the wording it sees first.

A benchmark is not authorization or clinical deployment

The April research did not test autonomous deployment and did not establish that o1 is safe for unsupervised diagnostic use. Its authors called for prospective trials—studies in which the system is evaluated during actual care with predefined safety and outcome measures.

That distinction also matters in regulation. The FDA’s current AI-enabled device list covers products that have met applicable premarket requirements for particular intended uses; the agency says it is still developing ways to identify foundation-model and large-language-model functionality more clearly in future updates.

A general reasoning model performing well in a journal experiment is therefore not equivalent to a medical device authorized for a defined indication. Deployment questions include which data the system may use, how clinicians review its output, how performance is monitored after implementation and who responds when the model and physician disagree.

The most credible role is a checked second opinion

The study supports evaluating o1-like systems as diagnostic support, especially when a record is long, fragmented or uncertain. A model could generate an additional differential, surface a possibility omitted from the initial assessment or identify conflicting details for a clinician to investigate.

That role preserves the study’s genuine value without assigning the model abilities the experiment never tested. A useful prospective trial would compare ordinary care with a clearly defined human-AI workflow, then measure more than agreement with the final diagnosis: missed emergencies, unnecessary testing, time to treatment, subgroup performance, clinician overrides and patient outcomes all matter.

The durable finding is not that an AI system can replace an emergency physician. It is that a reasoning model extracted diagnostically useful patterns from difficult hospital records and, on one small comparison, surpassed two physicians’ written answers. The newer multimodal evidence explains why the next step must be supervised clinical testing rather than autonomous use.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0