Voices

AI Chatbots: Helpful or Harmful in Health Advice?

AI chatbots like GPT-4o, Meta’s Llama 3, and Command R+ can ace medical licensing exams, correctly identifying textbook conditions in controlled lab tests at rates as high as 95 percent. But a study out of the University of Oxford, published in Nature Medicine, found that this knowledge doesn’t reliably translate when real people turn to these tools with real symptoms.

What the study found

Researchers led by Adam Mahdi at Oxford’s Reasoning with Machines Lab ran an experiment with nearly 1,300 participants. They gave half of the group medical scenarios, including a severe headache and new-mother exhaustion, and asked them to use an AI chatbot to figure out what was wrong and what to do next. The other half used traditional web search instead.

The results were sobering. When researchers tested the chatbots directly against the scenarios in the lab, the bots identified the correct condition with about 95 percent accuracy. But when human participants conversed with the same chatbots to work through their symptoms, accuracy dropped to under 35 percent for diagnosis and roughly 44 percent for recommending the right next step, such as whether to call a doctor or go to the ER.

Some individual scenarios show just how high the stakes can get. In one case, two participants described the same underlying condition in slightly different ways. The chatbot told one to go to the ER immediately; researchers later confirmed the condition was in fact life-threatening.

Why the communication gap happens

The core problem isn’t that the models lack medical knowledge. It’s how humans and chatbots talk past each other. Mahdi described the disconnect simply: the AI has the knowledge, but people struggle to extract useful advice from it. Study co-author Andrew Bean of Oxford made a similar point, noting that people frequently don’t know what information a model actually needs from them.

Part of the issue lies with users, who often share symptoms gradually, leave out relevant details, or phrase the same problem differently from one conversation to the next. But part of it lies with the chatbots themselves: unlike a physician in an exam room, they generally don’t push back, ask probing follow-up questions, or notice what a person is avoiding saying. Instead, they tend to respond to whatever the user gives them, listing possible conditions without helping the user work out which one actually fits.

A separate Mount Sinai-linked analysis raised a related concern. Even when a chatbot landed on the right diagnosis, it didn’t always convey the right sense of urgency. In more than half of emergency scenarios tested, the bots “under-triaged” cases, treating a condition as less serious than it actually was. In one instance, a chatbot failed to direct a hypothetical patient with a life-threatening case of diabetic ketoacidosis to seek emergency care.

Why this matters now

This isn’t a niche problem. Researchers estimate that roughly one in six American adults already consult an AI chatbot for medical information at least once a month, and separate polling found that more than a third of UK residents have used AI chatbots for mental health support. As people embed these tools more deeply into daily life, the researchers argue, the consequences of miscommunication scale right along with adoption.

Some worst-case outcomes have already surfaced publicly. Scientific American reported that a 75-year-old man in Seattle died of a treatable form of leukemia in December after he reportedly declined treatment based on inaccurate AI-generated information suggesting he had a rare complication.

A more complicated picture than “good” or “bad”

It would oversimplify things to call chatbots simply unreliable. In other controlled research, large language models have matched or even outperformed physicians on certain diagnostic reasoning tasks, and both OpenAI and Anthropic have since released dedicated health-focused chatbot versions to address these accuracy gaps. Physicians who commented on the findings, including UCSF’s Robert Wachter, describe the broader trend of patients arriving at appointments having already consulted a chatbot as “generally healthy,” even when it costs him time correcting bad advice during a short appointment.

The structural realities of healthcare access also favor AI in some cases. For people who are uninsured or face long waits for an appointment, a chatbot may offer the only accessible source of medical guidance available at all, even an imperfect one.

READ: What Is Pax Silica and Why Is Everyone Talking About It?

The bottom line

The Oxford findings, along with related work published in Nature Medicine and elsewhere, suggest that the underlying models’ medical knowledge isn’t the limiting factor in AI health advice right now. The conversation itself is.

Until chatbots get better at asking the right follow-up questions, and users get better at communicating what actually matters, the researchers conclude that none of the tested models are ready to replace professional medical care.

For now, experts recommend that people treat chatbot output as a starting point for a conversation with a real clinician, not a replacement for one.

Show More

Leave a Reply

Your email address will not be published. Required fields are marked *