Chapter 3 of 4
Six Reasons We Are Not Ready
Saying wait has a cost. People in counties without a single primary care doctor are not a hypothetical. But going now has a cost too — and the costs land on the same people. Here are the six things I keep coming back to. Tests are not patients. The studies are about AI-helped doctors, not AI alone. AI fails confidently. People are not ready. The first deployments will land on those who already get less. And the rules to hold anyone responsible do not exist yet.

Tap the picture to see it full size
I want to be careful here. Saying wait has a real cost. People in counties without a doctor are not a hypothetical. But going now has a cost too. And the costs land on the same patients.
These are the six things I keep coming back to.
1. Tests are not patients
Even the JAMA authors say this in passing. Their own sentence:
Performance on written examinations provides important but incomplete evidence of clinical competence.
That is from the paper itself (Bergman et al., 2026). Passing the doctor's exam is not the same thing as taking care of a patient. The exam asks for the textbook answer. The patient asks for the answer for them — sitting in this room, with this history, with the thing they are too scared to say out loud.
Think of a kid who can score perfectly on a math test about swimming. That does not mean you let them lifeguard the pool.
2. Most of the evidence is for AI-helped doctors, not AI alone
Read the studies the paper cites and a pattern shows up. The Kenya study was about doctors using AI tools (Korom et al., 2025). The Pakistan trial was about doctors trained to use language models well (Qazi et al., 2026). Even the NOHARM study compared the AI to physicians on benchmarks, not in the actual job of seeing real patients on its own (Wu et al., 2025).
All of that evidence supports the first paper's idea — that AI helps doctors do their job better. None of it tells us much about AI on its own, with no doctor in the loop, taking care of real people.
That is a different question. We barely have data for it.
3. AI fails confidently
Current AI systems have a particular kind of mistake that is dangerous in medicine. They get things wrong while sounding right. The technical word is hallucination — when the AI invents a fact and presents it as true.
Imagine a friend who never says I do not know. Every question gets an answer. The answers sound good. About 9 out of 10 times the answer is right. The other time, it is wrong, and the friend is just as sure as before.
With a doctor in the loop, this is recoverable. The doctor catches the wrong answer. Without a doctor in the loop, the wrong answer is the diagnosis. There is no second pair of eyes.
4. People are not ready
This is the reason most of the policy papers leave out, and the one I keep returning to.
Even if AI passed every exam, even if it had perfect data, even if it never hallucinated — people are not ready to be cared for by an AI alone when they are sick or scared.
A 2026 study in JAMA Network Open surveyed 3,000 US adults on what they actually want from medical AI (Bracic et al., 2026). The pattern was clear. People were much more comfortable with AI when a clinician was present. When the AI was approved by a trusted body. When the data behind it looked like them. The thing patients wanted most was a real human in the room.
A separate 2025 randomized survey of 1,762 US adults found that just mentioning a doctor used AI made some patients trust the doctor less (Chen and Cui, 2025). That is not a tech problem. That is a relationship problem.
Care is not only about the right answer. It is about being heard. About someone making eye contact when the news is hard. About the small human things that happen between two people in a room. We do not yet know how to give that without a person there.
5. The first deployments will land on people who already get less
Read the paper carefully and you will find the example the authors give. Autonomous AI is most needed, they say, in rural federally qualified health centers — clinics that serve communities without much else (Bergman et al., 2026). That is the example.
That is also the worry.
People in well-resourced places will keep their human doctors. People in poor and rural places will be the first to be cared for by an AI on its own. The two-tier risk is not theoretical. It is in the proposal, named clearly.
If autonomous AI works, this is a real expansion of access. If it does not work, the harms land on the people who already have the fewest options.
6. The rules to hold anyone responsible do not exist yet
Imagine an autonomous AI gives a patient the wrong diagnosis tomorrow and the patient is hurt. Who is responsible?
The paper says the developer is primarily responsible. The hospital that deployed it is secondarily responsible. That sounds reasonable. But the legal system to actually hold those parties to account does not exist yet. There is no AI medical malpractice law in any state. There is no precedent. There is no clear way for a patient to sue.
When a self-driving car hurts someone today, the lawsuit takes years and goes through layers of liability that nobody planned for. We are about to do the same thing in medicine, only with patients who cannot wait years.
The licensing idea is good. The order is wrong. We should build the rules first — then issue the licenses.
Where this leaves us
Six reasons. Some technical, some human, some legal. Together they are why I think the world is not ready for autonomous AI doctors today, even though the case for them is real.
The next chapter is about what we should be doing instead.
References
Bergman, A., Wachter, R. M., Emanuel, E. J. (2026). A Licensure Framework for Autonomous Clinical AI. JAMA. doi:10.1001/jama.2026.5483
Bracic, A., Spector-Bagdady, K., Towle, S., Zhang, R., James, C. A., Price, W. N. II. (2026). Factors for Patient Trust and Acceptance of Medical Artificial Intelligence. JAMA Network Open, 9(3), e260815. doi:10.1001/jamanetworkopen.2026.0815
Chen, C., Cui, Z. (2025). Impact of AI-Assisted Diagnosis on American Patients' Trust in and Intention to Seek Help From Health Care Professionals. Journal of Medical Internet Research, 27, e66083. doi:10.2196/66083
Korom, R., Kiptinness, S., Adan, N., et al. (2025). AI-based clinical decision support for primary care: a real-world study. arXiv. doi:10.48550/arXiv.2507.16947
Qazi, I. A., Ali, A., Khawaja, A. U., et al. (2026). Large language model diagnostic assistance for physicians in a lower-middle-income country: a randomized controlled trial. Nature Health, 1(2), 198–205. doi:10.1038/s44360-025-00007-8
Wu, D., Haredasht, F. N., Maharaj, S. K., et al. (2025). First, do NOHARM: towards clinically safe large language models. arXiv. doi:10.48550/arXiv.2512.01241
Written by
HurrozMore by Hurroz
Short link: hurroz.com/c/01FHB