I do not want my doctor replaced by a chatbot, but that doesn’t mean I don’t want them using it.

I absolutely want my doctor using the tool that might notice the drug interaction, rare diagnosis, missing test, or red flag that a tired human missed at 4:40 p.m.

That is the uncomfortable middle ground medicine is entering now. AI is not ready to practice medicine by itself. But it’s also getting too useful for serious doctors to ignore.

On Monday, Derya Unutmaz, an immunologist and professor, made the sharper version of that argument on X: GPT-5.5 Pro, he argued, is already better than almost all doctors, and soon not using AI in diagnosis and treatment should be considered malpractice.

I think he may be practically right before he is legally right.

That distinction matters. A patient can reasonably feel harmed by a doctor who refused to use a good AI second opinion long before a court says that refusal meets the legal standard for malpractice.

That gap between what patients expect, what doctors can defend, and what the law can enforce is where the next fight over AI in healthcare is going to live.

The BIG Story: Your Doctor Is About To Need An AI Second Opinion

The fastest way to misunderstand this story is to turn it into "AI versus doctors."

That is the internet version. It is also the least useful version.

The question is when does AI becomes part of what a competent physician is expected to use?

Imagine a doctor who refused to look at an X-ray or check drug interactions using a database.

I don’t mean because the machine was broken, or the image was unclear. Just because they did not trust the technology and preferred to go by instinct.

That doctor would not at all sound admirably human. They would sound reckless.

AI is not there yet. A general-purpose chatbot is not the same thing as an X-ray, a CT scan, a pulse oximeter, or a drug-interaction database. Those tools earned their place through evidence, workflow, regulation, training, and repeated use in real clinical settings.

But for some narrow parts of medicine, AI is already starting down that road.

Derya's claim landed because it says out loud what a lot of patients are going to start asking quietly: if a tool exists that can help catch what humans miss, why wouldn't my doctor use it?

My own reaction is pretty simple. I want my physician using every serious tool available if it leads to better answers and better care. I do not want magical thinking. I do not want a doctor copying and pasting my symptoms into a chatbot and calling it medicine. But I also do not want a physician pretending the last few years of AI progress did not happen.

The evidence is getting harder to shrug off.

And it is not just coming from model benchmarks.

A new Nature Health paper from Microsoft AI researchers, including Mustafa Suleyman, analyzed more than 500,000 de-identified health-related conversations with Microsoft Copilot from January 2026. The striking part is how personal common that usage already is: nearly one in five conversations involved symptom assessment or condition discussion, personal health queries rose in the evening and at night, and one in seven personal health queries concerned someone other than the user, such as a child, partner, or aging parent.

That matters because patients are not waiting for the medical system to decide whether AI belongs in healthcare. They are already using it as the first nervous stop between "something feels wrong" and "do I need to call someone?" We’ve already answered the first question: "will patients use AI?"

The new one, however, is equally important: "Will doctors and health systems meet them there responsibly?"

OpenAI says its GPT-5.5 Instant health improvements made the model better at recognizing when urgent care may be needed, asking for more context, explaining uncertainty, and producing useful health responses. In one physician-reviewed evaluation, OpenAI said GPT-5.5 Instant responses were rated higher than physician-written responses across several criteria.

That is not independent proof that ChatGPT should practice medicine. It is a company-run evaluation, and it should be read with the usual caution. But it is still a signal.

The academic evidence is pointing in the same direction, with more caveats attached. Stanford Medicine summarized a study showing that doctors paired with a chatbot performed better on nuanced clinical decisions than doctors using conventional resources. A Stanford-Harvard State of Clinical AI summary found the most consistent benefits when AI supports clinicians rather than replaces them.

And a 2026 study published in Science found that an OpenAI o1-series model outperformed physicians on several text-based clinical reasoning tasks, including work based on real emergency department cases. The authors were careful about the limits: this was not live patient care, and it does not mean doctors can be removed from the process.

That caveat is the whole story.

Medical care goes far beyond simple diagnosis. It is the exam room. It is the patient who forgets to mention the relevant symptom until the end of the visit. It is the family history, the local hospital system, the insurance constraint, the medication list that is almost but not quite accurate, the trust required to ask the embarrassing question, and the judgment to know when a model is being confidently wrong.

When the patient is fighting an insurance system that doesn’t want to treat them AND doesn’t want to pay their doctor, this presents a unique opportunity for the future.

This is where Marc Andreessen's comments in the New York Post are useful. His optimism about AI in healthcare is not hard to understand. No individual doctor can read every new paper, track every medication interaction, and remember every obscure presentation of every disease. AI can help with exactly the kind of pattern-matching and literature-scale reasoning that humans are structurally bad at.

That is especially obvious for elderly patients on many medications, patients with rare conditions, and cases where the first answer is plausible but wrong.

The missing layer is implementation.

It is one thing for AI to be smart. It is another thing for a health system to use it safely.

Who chooses the tool? Who checks whether it works for this patient population? What happens when the model is wrong? Is the patient told AI was used? Does the doctor document the AI output? Is the tool FDA-cleared for this use, or is it a general assistant sitting outside the regulated clinical workflow? What data leaves the clinic? Who is liable if the doctor follows the tool and harms the patient? Who is liable if the doctor ignores it and misses something obvious?

That last question is the one coming for medicine.

Legal malpractice is a higher bar than "this seems dumb in hindsight." Courts usually ask whether a clinician acted like a reasonably competent clinician would have acted under similar circumstances. That standard changes slowly. It moves through evidence, professional guidelines, institutional protocols, insurance norms, documentation habits, and actual cases.

So no, a viral benchmark does not instantly make AI a legal requirement.

But practical malpractice moves faster than legal malpractice.

If a patient has a bad outcome because a doctor missed a drug interaction that an AI system would have flagged instantly, that patient is not going to comfort themselves with the phrase "evolving standard of care." They are going to ask why the doctor did not check.

That is why Derya's claim is worth taking seriously even if the timeline is aggressive.

He may be early on the legal standard. He may be right on the practical expectation.

What Comes Next…

This will probably not arrive as one big dramatic mandate.

It will happen in layers.

First, AI becomes administrative. That is already underway. Ambient scribes listen to visits, draft notes, summarize encounters, and reduce some of the paperwork that has made modern medicine miserable for both doctors and patients. This is the beachhead because it solves an obvious pain point without asking AI to make the final clinical call.

Then AI becomes a second opinion.

Not "What disease does this patient have?" but "What else should we consider?" Not "Write the treatment plan," but "Check these medications for interactions." Not "Replace the doctor," but "Challenge the first answer."

That is where the near-term case is strongest. AI is very good at generating differentials, surfacing possibilities, translating medical complexity, and catching patterns across large amounts of text. Used well, it can make a good doctor sharper. Used badly, it can make a rushed doctor more confident in a mistake.

After that comes institutional protocol.

Hospitals and practices will start requiring AI checks in defined workflows: medication reconciliation, radiology support, emergency triage, discharge instructions, complex differentials, prior authorization documentation, and maybe rare-disease review.

That is when the malpractice clock really starts.

The key moment is not "AI exists." It is when failing to consult a validated tool in a defined clinical situation starts to look outside normal professional practice.

The awkward part is that doctors may get squeezed from both sides.

If they rely too much on AI and it is wrong, they may be blamed for outsourcing judgment. If they refuse to use AI and miss something the system could have caught, they may be blamed for ignoring available support.

That is not fair if hospitals, vendors, regulators, and insurers leave physicians holding the whole bag. The American Medical Association has already flagged physician liability as one of the central issues in healthcare AI. The FDA keeps a list of AI-enabled medical devices, while groups like the Coalition for Health AI are pushing model-card-style transparency for healthcare AI systems through efforts like the CHAI Applied Model Card.

That is the boring infrastructure that makes the flashy claim real.

Not the tweet. Not the demo. Not the benchmark.

The forms. The audits. The workflows. The training. The logs. The consent language. The answer to "who is responsible when this goes wrong?"

The best version of the future is not AI practicing medicine alone. It is a triad: patient, clinician, and AI system. The patient brings context and values. The clinician brings judgment, examination, accountability, and care. The AI brings memory, pattern recognition, literature-scale search, and a willingness to ask the annoying extra question forever.

That last part matters.

Good doctors already do this in human form. They ask one more question. They consider the unlikely thing. They check the interaction. They pause before assuming the obvious diagnosis is right.

AI will not make bad medicine good by magic.

But if it can make that kind of checking cheaper, faster, and more consistent, then eventually patients will expect it.

And once patients expect it, professional norms will start to move.

And once professional norms move, the law eventually follows.

Corey’s Field Notes

Builders Have To Explain Themselves

One of Andreessen's better points in the Post interview had nothing to do with healthcare: if builders want to change the world, they have to explain what they are building before someone else explains it for them. That is not just a PR lesson. It is now a survival skill. AI companies keep acting surprised when the public, regulators, workers, and institutions fill the silence with fear. If your product touches labor, learning, medicine, media, warfare, or childhood, "trust us, it is complicated" is not a communications strategy. It is an invitation.

The Best AI Workflows Are Getting Less Dramatic

The most useful AI habits I keep seeing are not cinematic. They are boring little systems: turn this meeting into decisions, turn these notes into a plan, check this memo for missing objections, summarize what changed since yesterday, compare these options against my actual constraints. That is not as exciting as an autonomous everything-agent. It is also where the real productivity gains are starting to show up. The magic is less "AI did my job" and more "AI removed the friction that kept me from doing the important part."

AI Tutoring Is Going To Pressure Schools From Below

Andreessen also pointed to AI tutors as endlessly patient, always-available help. The institutional education debate will take years. The household version is already simpler: students can get explanations at midnight, parents can get unstuck, and adults can teach themselves skills without waiting for a course to exist. The risk is that quality becomes uneven and dependence grows quietly. But the pressure is obvious. Once people get used to patient, personalized help on demand, normal one-size-fits-all instruction starts to feel much older very quickly.

One Useful Thing

Bring AI into your next medical appointment without turning into That Patient.

The move is not to announce that you have solved medicine. Please do not do that. The move is to use AI to become clearer, better organized, and harder to accidentally dismiss.

Before the appointment, prepare four things:

  • A short symptom timeline.

  • Your current medications and supplements.

  • Any allergies, prior diagnoses, recent tests, or relevant family history.

  • Three questions you want answered before you leave.

Then ask your doctor something like:

"Would it be useful to run this through a clinical decision-support or AI tool to check for medication interactions, red flags, or diagnoses we should rule out?"

That phrasing matters. You are not asking the AI to replace the physician. You are asking whether another layer of checking would improve the visit.

A few good follow-ups:

  • "What diagnoses are you considering?"

  • "What have we ruled out?"

  • "What symptoms would make this urgent?"

  • "Are there medication interactions or contraindications we should double-check?"

  • "If this does not improve, what is the next step?"

The goal is not to annoy your doctor and make things awkward. Just explain you’re invested in your care and want to better understand the full picture of your health.

Either way, you’ll find out if the doctor is the right fit for you.

What I’m Watching This Week

GPT-5.6 Watch

OpenAI has already previewed GPT-5.6, but the broader rollout is still the thing to watch. The credible reporting says the U.S. government asked OpenAI to limit early access to a small group of trusted partners; the rumor layer is that broader availability could come soon if the review process clears. Translation: the model story is now also a governance story.

Treasury’s AI Bubble Report

NOTUS reports that a draft Treasury Department analysis warns AI could become a financial-system risk if valuations, data-center financing, productivity expectations, or infrastructure bottlenecks break the wrong way. Treasury is publicly distancing itself from the draft, so I’m treating this as leaked/draft-policy signal, not settled government position. Still, it matters that AI is now being analyzed as a possible systemic finance story, not just a tech story.

Meta’s Watermelon

The rumor to watch: Meta’s next model, codenamed Watermelon, has reportedly caught up with OpenAI’s GPT-5.5 on internal benchmarks. That’s internal, undisclosed, and very much not the same as public proof. But if Meta is even close, the frontier race gets messier fast: model quality, compute spend, talent raids, and distribution all start colliding.

/

Reply

Avatar

or to participate

Keep Reading