The credential was doing several jobs at once
For a long time, we attached trust to the professional because the professional was where the scarce assets lived.
The doctor had the medical knowledge. The lawyer had the cases. The accountant had the tax code. The engineer had the calculation methods. The professional had access to the records, the language of the field, the accepted standards, and the judgment that came from experience. The credential became a reasonable proxy for the whole bundle.
That bundle is beginning to come apart.
Knowledge is no longer confined to the professional’s memory or bookshelf. A capable system can search, retrieve, compare, summarize, calculate, challenge, and keep working long after the appointment clock has expired. The individual may possess the deepest longitudinal context because the individual has lived it. The machine may be better at holding thousands of details in view. The professional still has examination skills, practical experience, institutional access, legal authority, and accountability.
Those are different capabilities. We keep talking as if they must all reside in one person.
They do not.
The credential is still real. It is no longer the whole system.
That distinction matters because the public argument keeps collapsing into a primitive contest: human versus machine. Would you trust a doctor or an AI? Would you let a robot make the decision? Do you want a person in the loop?
Those questions sound sensible because they are familiar. They are also badly formed.
A person with a medical degree can make an error. A language model can make an error. A hurried clinician using a poorly integrated model can make a new kind of error that neither would have made alone. A disciplined patient using longitudinal records, authoritative sources, explicit uncertainty, and a clear escalation boundary may outperform an ordinary fifteen-minute encounter on a bounded research task. That same patient and system may be dangerously unqualified to decide whether chest pain can wait until morning.
The useful comparison is not human against machine. It is system against system, task against task, under a defined set of conditions.
The evidence is already uncomfortable
The evidence does not support the cartoon version in which AI is either an infallible synthetic physician or a stochastic parlor trick.
In a randomized clinical trial published in JAMA Network Open, physicians using GPT-4 did not significantly outperform physicians using conventional resources on diagnostic reasoning. GPT-4 alone, however, scored higher than the conventional-resource physician group. The investigators did not conclude that doctors were obsolete. They pointed to something more operationally important: giving a capable tool to a professional does not guarantee that the combined system will use it well. The interface, workflow, trust calibration, and operator behavior all matter.[^1]
Google’s AMIE research makes the other side difficult to dismiss. In a 2025 Nature study using simulated text consultations, the system was rated higher than primary-care physicians on most of the measured specialist and patient-centered axes and produced stronger differential diagnoses. A later multimodal study involving dermatology images, ECGs, and clinical documents reported similarly strong results across simulated telehealth consultations.[^2][^3]
Those were controlled simulations, not permission to replace ordinary clinical practice with a browser window. But “it was only a benchmark” is not an intellectually serious way to ignore repeated performance on the cognitive work the benchmark was designed to measure.
The negative evidence is just as important. A 2026 JAMA Network Open evaluation found that frontier models performed poorly when asked to build early differential diagnoses from sparse initial presentations. The authors concluded that the systems they tested were unsafe for unsupervised patient-facing use in that setting.[^4] In a randomized study of clinicians, standard AI assistance improved diagnostic accuracy, while systematically biased AI reduced it. Explanations did not reliably rescue the clinicians from the bad model.[^5]
A 2026 consumer study in dermatology found another boundary worth paying attention to. AI assistance improved people’s ability to name likely skin conditions, but it did not improve the accuracy of what they should do next.[^6]
That is nearly the whole thesis in one result.
Inference is not action. Naming is not treating. Analysis is not authority.
The machine may be the best component in the system for one stage and the wrong component for the next.
“Keep a human in the loop” is not a safety system
When people sense the stakes of AI, they often retreat to a reassuring phrase: keep a human in the loop.
Human governance does not require human dominance of every task.
Which human? In what position? With what information? Under what time pressure? With the authority to do what? Is the human genuinely evaluating the machine, or decorating its answer with a credential?
Human involvement can improve a system. It can also make the system worse.
If the AI is better at a bounded task and the human overrides it for weak reasons, the human has degraded the result. If the AI is confidently wrong and the human defers because the explanation sounds polished, the human has degraded the result. If both parties share the same missing context, adding one to the other has not created independence. It has created agreement.
A 2026 JAMA Viewpoint synthesized this emerging and inconvenient evidence. Its argument was not that autonomous AI should take over medicine tomorrow. It was that the comfortable assumption that an AI-assisted physician must always be superior to AI alone has not been proven. Across experiments summarized by the authors, hybrid performance could fall below the stronger component, especially when people failed to recognize whether the task sat inside or outside the machine’s competence.[^7]
This pattern is not confined to medicine. Research involving hundreds of consultants found what the authors called a “jagged technological frontier.” On some tasks, AI substantially improved performance. On tasks outside its frontier, reliance on AI made performance worse. The advantage belonged not merely to people who had the tool, but to people who could navigate that uneven boundary.[^8]
Boundary recognition is the skill we are actually going to need.
Not faith in machines. Not reflexive defense of humans. Boundary recognition.
The skilled operator changes the equation
I would not tell the general public to rely on a consumer chatbot for medical advice without qualification. That would be reckless.
Many people do not know how to inspect sources, distinguish a confident sentence from a supported conclusion, test an assumption, protect sensitive data, or recognize when a question has crossed from research into diagnosis or treatment. That is not an insult. Most of us were never taught to do any of it. The products themselves frequently make unsupported confidence feel effortless.
But the phrase “the general public” can also conceal an important fact: users are not interchangeable.
A person who can supply clean longitudinal data, separate observation from interpretation, constrain the question, demand primary evidence, challenge the output, test for contradiction, record uncertainty, and escalate at the right moment is not operating the same system as someone typing a symptom into a generic chat window at two in the morning.
A chatbot is not a clinical system any more than a search box is a physician.
The system includes the model, the evidence, the data, the prompt or procedure, the person operating it, the verification method, the decision boundary, and the escalation path. Change any one of those and you may change the safety and quality of the result.
That is why the statement “AI gives medical advice” is almost useless. Which AI? Using what records? Grounded in which sources? Asked to do what? Verified by whom? Connected to what authority? Allowed to trigger which action?
The same questions should be asked of the professional encounter.
Which professional? With how much time? Looking at which records? Working from memory or current evidence? Seeing the whole history or the last note? Coordinating with whom? Accountable for which part?
Once both sides are described honestly, the inherited hierarchy becomes less stable.
For a growing class of bounded cognitive tasks, a capable person with a governed AI system may already receive more analytical value than the ordinary professional encounter provides.
That is the claim.
It is deliberately narrower than “AI is a better doctor.” It is also much harder to wave away.
Trust should move with the work
The right operating model is not to hand everything to the machine or reserve everything for the professional. It is to assign each layer of work to the component best qualified to perform it, then make the transfers explicit.
The stack looks something like this:
Today we often pretend the credential answers all ten questions. It does not. AI enthusiasm makes the equal and opposite mistake of pretending a high-performing model answers all ten. It does not either.
In my kidney-stone research, AI could help organize observations, find patterns in data, preserve longitudinal context, generate hypotheses, and retrieve evidence. I could judge whether the analysis reflected my actual history and use it to formulate better questions. My doctor still had access to physical examination, ordering, prescribing, clinical accountability, and the ability to act within a licensed system.
The failure was not that a doctor remained involved. The failure was that the conventional encounter had no reliable mechanism for integrating the better analysis that arrived with the patient.
The design is broken.
Waymo and the emotional arithmetic of error
Autonomous driving exposes the same tension because the mistakes are visible and frightening.
When a human driver causes a crash, we tend to process it as a familiar tragedy. When an autonomous vehicle causes one, we process it as evidence about the legitimacy of the entire category. The machine is asked to be perfect before it is permitted to be better.
That is not a rational safety standard.
A peer-reviewed 2025 analysis of 56.7 million rider-only Waymo miles reported significantly lower crash rates than matched human benchmarks for injury-reported, airbag-deployment, and serious-injury crashes. It also found large reductions in several intersection categories and no statistically significant disbenefit among the crash groups studied. The authors were affiliated with Waymo, which belongs in the evidence record, and continued independent evaluation is essential.[^9]
The relevant question is not whether an autonomous system can make a mistake. It can. The question is which system produces fewer consequential mistakes within the same operating domain, how we know, and what happens when the system reaches its boundary.
Medicine will face the same emotional arithmetic.
If an AI misses a diagnosis, the miss will be offered as proof that machines cannot be trusted. If a human misses the same diagnosis during a rushed visit, we may treat it as an unfortunate feature of a complicated profession. Neither response is good enough.
We should be harder on both systems and more honest about the comparison.
This does not end with doctors
I believe the same disaggregation will reach lawyers, accountants, engineers, project managers, consultants, and nearly every other profession built partly on scarce access to organized knowledge. Medicine gives us an evidence-rich proving ground. It does not, by itself, prove what will happen in every other field.
That does not mean every task becomes a prompt. A lawyer still exercises legal judgment, represents a client, negotiates, appears in court, and carries duties that a model cannot assume. An accountant signs work, interprets facts, manages controls, and owns professional obligations. An engineer works inside physical, regulatory, and safety constraints. A project manager coordinates people, conflict, timing, risk, and commitment.
But large pieces of the cognitive work inside those professions can now be decomposed. Research can be performed separately from judgment. Pattern detection can be separated from authority. Drafting can be separated from acceptance. Calculation can be separated from responsibility. Monitoring can be separated from intervention.
Once that happens, “Are you a professional?” is no longer enough to tell us whether a particular piece of work will be done well.
We will need to ask who or what has the best context, which component is strongest at the task, how the result was verified, and where authority should return to a human.
This is not the elimination of professionalism. It is the end of professional monopoly over every layer of the work.
The next divide is competence, not access alone
People are already using chatbots for health questions. In a 2026 Pew Research Center survey, roughly a third of U.S. adults said they had used one for at least one of several health-related purposes. People with more education and higher incomes were considerably more likely to do so. Another Pew survey found that many Americans struggle to judge whether health information is accurate or to resolve conflicts between sources.[^10][^11]
That creates two risks at the same time.
The first is obvious: people may trust weak systems too much.
The second receives less attention: people who learn to build and govern strong systems may pull away from people who never acquire those skills or cannot access the necessary models, records, devices, and evidence.
AI literacy is not going to mean knowing how to make a chatbot write an email. It will mean knowing how to assemble a trustworthy process around a fallible machine.
It will include source judgment, data hygiene, uncertainty, privacy, comparison, adversarial testing, decision rights, and escalation. In health contexts, privacy deserves special emphasis. Consumer health tools do not all sit inside the same protections people associate with a hospital or physician, and both regulators and the World Health Organization have warned about privacy, bias, security, and automation risks.[^12][^13]
The people who understand those boundaries will not merely get faster answers. They may ask better questions, detect problems earlier, enter professional encounters better prepared, and refuse weak reasoning regardless of whether it comes from a person or a machine.
That is a meaningful advantage. Left alone, it will become another form of inequality.
The standard should be earned authority
We are going to trust machines with consequential work. In some domains, we already do. The choice is not between a permanently human world and a recklessly automated one.
The choice is whether we build systems that make competence, evidence, authority, and escalation visible, or continue relying on proxies that technology has already begun to dissolve.
A good system should be able to tell us what it observed, what it inferred, what evidence it used, where uncertainty remains, what it is authorized to do, and when it must stop. A good human operator should be able to interrogate those same layers. A good professional should be willing to incorporate better analysis without treating it as a threat to the credential.
The patient bringing a stronger analytical system into the room is not an attack on medicine. It is an opportunity to make the whole system better.
But that opportunity requires giving up a comforting fiction. A human face is not proof of sound judgment. A credential is not proof of current knowledge. A fluent machine is not proof of truth. And adding a human to an AI does not automatically produce the best of both.
Trust should move with demonstrated capability. Authority should move more carefully. Action should remain constrained by consequence. Escalation is not failure. It is part of competence.
Trust does not belong permanently to the face in front of us or the voice in the box. It belongs, temporarily and conditionally, to the system that has earned it.