In Harvard study, AI offered more accurate emergency room diagnoses than two human doctors
A team from Harvard Medical School and Beth Israel Deaconess Medical Center conducted a new study evaluating the effectiveness of AI in real-life medical situations. The research, published in Science, compared OpenAI’s language models with internal medicine attending physicians using cases from the emergency department.
Study methodology and key findings
Researchers reviewed 76 patient cases from the Beth Israel emergency room. They compared the diagnoses made by two doctors to those suggested by OpenAI’s o1 and 4o models. The AI and human diagnoses were then reviewed by two other physicians who were unaware of their source.
The results revealed that the o1 AI model matched or exceeded the accuracy of the two attendings, particularly during the first triage phase. Notably, o1’s diagnoses were exact or very close to the final outcome 67% of the time, outperforming the two human doctors, who achieved 55% and 50% respectively.
Important caveats and expert perspectives
The study’s authors caution that their work doesn’t mean AI is ready to make real-time, high-stakes clinical decisions. Instead, they recommend further clinical testing of these technologies in actual hospital environments. They also note that the AI’s evaluations were based solely on written medical records, not physical exams or other types of data.
Some experts highlight that these comparisons were made with internal medicine doctors, not emergency specialists, and stress that human connection is crucial for critical patient care and decision-making. Dr. Kristen Panthagani, an emergency physician, argued that assessing immediate risk is the primary goal in the ER, more so than pinpointing a final diagnosis.
For more details and perspectives, read the original article by Anthony Ha at TechCrunch.