An opinion piece in the medical journal JAMA predicts that autonomous AI will outperform any doctor-AI pairing at medical reasoning. The authors want to stop regulators from writing a human in the loop into the rules.
Knowing who wrote the JAMA piece helps explain its angle: Lead author Ezekiel Emanuel is a bioethicist at the University of Pennsylvania and one of the architects of Obama’s healthcare reform, making him a heavyweight in US health policy. Coauthor Neal Khosla is CEO of the AI telemedicine company Curai Health, and his father, Vinod Khosla, is an investor in both OpenAI and Curai Health. Two of the authors stand to gain directly from the future the piece champions.
That future clashes with the position held by physician groups like the American Medical Association and the American College of Physicians, which say AI should support doctors, never replace them. Medical professor Robert Wachter goes further in his book, calling AI-only care the “economy class” of medicine. The JAMA authors consider this ranking unproven and offer two arguments against it.
The first concerns the research. In most studies since 2024, AI on its own matches or beats doctors at the five core reasoning tasks of medicine, namely taking a patient history, making diagnoses, choosing tests, treating according to guidelines, and managing chronic disease.
Google’s conversational system AMIE scored higher than primary care doctors in almost every category during simulated patient conversations. Across 377 complex cases, ChatGPT o3 named the correct diagnosis first 60 percent of the time, compared with 15.9 percent for 20 internists. Microsoft’s diagnostic orchestrator found the correct diagnosis under budget constraints about four times as often as doctors, at lower cost. The authors dismiss most studies with the opposite finding as outdated or methodologically weak, for example because they left out the best models.
The second argument is a forecast about where things are headed, and it forms the heart of the piece. The models are improving fast, while doctors lose their own skills through AI use, as a Lancet study on colonoscopies suggests. The gap should only widen from here.
Once the machine is clearly ahead, the doctor doing the checking turns from a safety net into a source of error. According to the authors, a meta-analysis of 106 experiments backs this up. When the human is better, the combination helps. When the AI is better, the human makes the result worse by overruling the system in the wrong places. In a study using real patient cases, GPT-4 alone scored 92 percent on diagnostic reasoning, while doctors with access to the same model scored just 76 percent.
The authors use chess as a historical parallel: After Deep Blue beat Kasparov in 1997, human-machine teams dominated for years, until AI began beating the teams too starting in 2017.
A warning to the rule writers
From this the piece draws its real message: Locking in guidelines that require a doctor to make the final call could cement a form of care that will soon fall behind. By 2030, the authors expect autonomous AI to be ready for some, maybe many, workflows, limited to cognitive tasks. Liability, payment, regulation, and medical training all need a rethink now, they argue.
There are however limitations to their message, as they admit. Almost all the evidence comes from simulations of single tasks, not real patient care, and the handoff of information between human and model counts as a weak spot.
Physical procedures like surgery, childbirth, and colonoscopies will also stay with humans for now, since the robotics aren’t up to it. And autonomous systems fail in ways doctors don’t, through hallucinations, internet outages, or cyberattacks. Those risks have to be weighed against the higher accuracy, the authors concede.




