Please don’t trust your chatbot for medical advice
Four separate studies all point in the same direction
Remember how I used to say that large language models are “frequently wrong, never in doubt”, and how I warned three years ago on 60 Minutes that they were purveyors of “authoritative bullshit” that should not be trusted?
That’s still true – and it very much applies in medicine.
And that matters, a lot. Because a large fraction of the population has begun to turn to chatbots for medical advice.
Two relevant new studies are reported today in the Washington Post, in a damning article.
The first new study, published by BMJ (affiliated with the British Medical Association) in a peer reviewed journal, and entitled “Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit”, studied five popular chatbots (Gemini, DeepSeek, Meta AI, ChatGPT and Grok), about one year ago, prompting each with 10 questions about things ranging from cancer to vaccines and nutrition, in open-ended dialogues, and reporting that nearly half of the responses were highly problematic. Worse, “chatbot outputs were consistently expressed with confidence and certainty”. The responses were also filled with hallucinations and fabricated citations.
All of this – the hallucinations, mistakes, and overconfidence – is entirely typical of LLMs, and entirely problematic in medicine. As the authors put it, in somewhat academic language, but entirely accurately, “continued deployment without public education and oversight risks amplifying misinformation.”
The second new study, published in JAMA Network Open, affiliated with the American Medical Association, called “Large Language Model Performance and Clinical Reasoning Tasks” looked at 21 frontier models across 29 questions, and reported that “despite progress, current LLMs remain limited in early diagnostic reasoning and cannot yet be relied on for unsupervised patient-facing clinical decision-making.”
And the Post article actually only reported part of the new scientific literature on LLMs and medicines. Two other new studies that they missed only add to the concerns.
One, published in Nature Medicine, was called “Reliability of LLMs as medical assistants for the general public: a randomized preregistered study”. This one focused on “whether LLMs can assist members of the public in identifying underlying conditions and choosing a course of action”. Again the results were both clear and troubling. LLMs “identified relevant conditions in fewer than 34.5% of cases… no better than [a] control group”. Here the problem wasn’t so much that the LLMs lacked access to proper information — the same study showed that the models could do better in the hands of trained physicians — but that patients don’t know how to guide the LLMs to the right places.
In a recurring theme, we see that LLMs don’t know what they don’t know; they work decently well with the information they’ve got but don’t know how to conduct clinical interviews, and in the hands of the lay public can easily give bad advice because the proper questions never get asked, either by the patient or the LLMs. (An expert doctor might use the LLM to better effect, by asking the right questions.)
Still another new study, also published recently in Nature Medicine, entitled ChatGPT Health performance in a structured test of triage recommendations, found that “Among gold-standard emergencies, the system undertriaged 52% of cases” and concluded that “These findings reveal missed high-risk emergencies and inconsistent activation of crisis safeguards, raising safety concerns that warrant prospective validation before consumer-scale deployment of artificial intelligence triage systems.”
As a scientist, I am always looking for converging evidence. Four studies in four journals published in the space of a few months reaching essentially the same conclusion is a crystal clear indicator that chatbots, especially when used by amateurs, simply cannot be trusted.
On a personal note, my friend Ben Riley lost his father recently, and Teddy Rosenbluth of The New York Times wrote a long, moving article about how his father was mislead by A.I regarding his leukemia.
I hope you will get a chance to read that, and also Ben’s own blog about the sad situation.
There will always be better models, but for now, and until proven otherwise, we should not take the apparent “confidence” of large language models — itself an illusion of how they are trained — to mean that we should trust large language with our lives.


Thank you for sharing this story, Gary, it's been restorative to me to have so many people learn a little about my dad. I know you've been grieving the loss of your mother recently as well.
It's been interesting to see the public response to Teddy's reporting on what happened. If you peruse the comments on the NYT, there are many doctors who express frustration at the unreliability of the tools. For example:
"As a primary care physician, I have used AI tools to explore possible causes of difficult to diagnose patient symptoms. While helpful in bringing forth diagnosis that might have been missed, as one drills down on how to manage a particular disease I have seen it confidently assert erroneous conclusions. When I point out the scientific inconsistencies of what it has asserted it reverts to sycophantic praise for my intelligence and abruptly changes its recommendations. What worries me is that I was only able to spot it's inaccuracies because of my depth of knowledge and experience in the subject, something that the average user does not have."
Likewise, people who've used LLMs to "self diagnose" their treatment report very mixed results. It's not that everyone has a bad experience, but when they do, the magnitude of the mistake is worrying. For example:
"I found real comfort and value in the AI responses. You can't just message your doctor at 11 p.m. and get an immediate answer as to why you're experiencing some new and confusing symptom. So I get how AI can become a trusted medical 'friend.' However, at the end of the day, AI was completely wrong about my diagnosis. In fact, out of probably 20 different possibilities it listed, my real condition was never an option it presented."
It is mystifying to me that we are conducting this mass experiment on society at scale. We don't allow anything like this with medical drugs, yet we seem to be ok with just watching to see what happens with AI.
And in my father's case, it brought him great pain, and death.
Do people who think that chatbots are conscious also believe that airplanes are birds?