Adam Rodman has vivid memories of the incident. He got stuck on a patient case while he was a second-year medical student. He went to the library, looked through the catalog, photocopied documents, and returned them to the team, just like medical students did in the early 2000s. Two hours were needed. Today, while still standing in the patient’s room, he takes out his phone and obtains the same level of detail in fifteen seconds.
This quiet, seemingly insignificant change is at the core of what Harvard researchers are now referring to as one of the most important advancements in clinical medicine in decades. Physicians and computer scientists from Harvard Medical School and Beth Israel Deaconess Medical Center conducted a significant new study that was published in the journal Science. They discovered that a large language model performed better than doctors on a number of fundamental clinical tasks. Independent experts are taking the results seriously because they were specific enough and the methodology was meticulous enough.
In the most well-known experiment, 76 actual patients showed up at an emergency room in Boston. The same electronic health records, including vital signs, basic demographics, and a few lines from a nurse’s intake notes, were given to an AI (OpenAI’s o1 reasoning model) and pairs of human doctors. In 67% of cases, the AI correctly or nearly correctly diagnosed the patient based on that same limited picture. Between fifty and fifty-five percent were managed by human doctors. The AI’s accuracy increased to 82 percent when more clinical information became available, while the experienced doctors’ accuracy was between 70 and 79 percent.
One instance is particularly noteworthy. A patient with worsening symptoms and a blood clot in the lungs showed up. The human physicians believed the blood thinners were not working. The AI also identified the patient’s history of lupus, indicating that the inflammation may be autoimmune rather than clot-related. It was correct.
Reading that makes it simple to envision emergency medicine undergoing some sort of algorithmic takeover. The researchers take care to refute that. “I don’t think our findings mean that AI replaces doctors,” stated Arjun Manrai, the head of an AI lab at Harvard Medical School and one of the study’s lead authors. According to him, this indicates that the technology has advanced to the point where it merits the same thorough clinical testing that is applied to any novel medical intervention: actual trials conducted in actual settings with actual accountability.
Practically speaking, the AI in this study was acting like a very thorough consultant going over a paper chart. It was unable to detect the patient’s distress, read their body language, or pick up on details that are not recorded in an electronic health record. There is no denying the importance of those gaps in clinical care. However, as Rodman stated, a new model of care is probably in store, one that is more akin to a triangle with the doctor, the patient, and an AI system operating in tandem rather than AI taking the place of the doctor.

AI is already being used by nearly 25% of American doctors to help with diagnosis. According to surveys conducted by the Royal College of Physicians in the UK, 16% of physicians use it on a daily basis. Clinicians experimented with prior authorization letters and administrative tasks before moving on to clinical reasoning, which contributed to the tools’ widespread use.
Isaac Kohane, chair of the Department of Biomedical Informatics at Harvard, remembers using a complicated pediatric endocrinology case to test an early iteration of GPT-4. It provided accurate answers to all the questions. “It was truly a quantum leap,” he stated, “from anything that anybody in computer science who was honest with themselves would have predicted in the next ten years.”
All of this has genuine worries running through it. Medical AI that has been trained on historical data inherits historical biases in terms of who was treated, how that treatment was administered, and whose symptoms were taken seriously. Artificial intelligence (AI) systems continue to have hallucinations and present fake facts with the same assurance as true ones. Furthermore, there is currently no official framework for accountability when AI leads to an incorrect diagnosis. Another issue brought up by Dr. Wei Xing of the University of Sheffield is that there is preliminary evidence that physicians may unintentionally accept the AI’s response rather than using their own judgment, a tendency that could gradually weaken clinical judgment.
Subtle risks also exist, such as the technology being implemented narrowly and optimized for administrative throughput or billing efficiency rather than the things that truly matter in a patient’s life. “If we don’t think big, if we don’t try to rethink how we’ve organized medicine, things might not change that much,” stated Rodman, who has spent years considering this as a researcher and clinician.
It’s difficult to avoid sitting with that tension. Harvard’s findings are truly startling. It’s not insignificant if an AI can identify a lupus connection in a hurried emergency room, catch what seventeen doctors missed, or process a treatment plan more accurately than the typical attending physician. However, there is still a significant gap between a compelling research finding and improved patient care in real hospitals, with all the mess that goes along with it. There are the tools. Now, the question is whether medicine will make good use of them.
