Skip to main content
doctor on a computer

ChatMD?

Issue 53: October 2026
21 min read
by James Krieger

Overview

  • What did they test? Researchers tested whether LLMs could assist people in identifying underlying medical conditions and choose a course of action in ten medical scenarios. Subjects were randomly assigned to receive medical assistance from an LLM or a source of their choice (control).
  • What did they find? When tested alone, LLMs completed scenarios accurately. However, when people used these LLMs for medical advice, the LLMs identified relevant conditions in less than a third of the cases and disposition in only 44% of the cases. Neither was better than the control group; in fact, LLMs were worse at identifying relevant conditions. 
  • What does it mean for you? You should not rely on LLMs like ChatGPT, Gemini, Claude, or CoPilot for medical advice. A basic web search will be more likely to give you accurate information. That said, nothing can replace the accuracy of getting advice from your doctor.

What’s the Problem?

Purpose

Artificial intelligence (AI), or really Large Language Models (LLMs)/chatbots seem to be all the rage these days. People have been using ChatGPT, Gemini, Claude, Copilot, or other LLMs for coding, for web searches and summaries (despite being wrong often), for image generation, etc. People are also using them for medical advice and diagnosis of medical conditions; there are anecdotes of people successfully using LLMs to diagnose their own conditions. In fact, one in six American adults are consulting AI chatbots for health information at least once per month 1.

But how good are LLMs for medical tasks? On one hand, LLM scores on medical knowledge benchmarks are similar to passing a U.S. Medical Licensing Exam 2. LLM-generated clinical documents are rated as equivalent or better to those written by doctors 3 4. However, performance in some controlled settings and benchmarks like this may not translate to real-life performance when people are using them for typical tasks. For example, radiologists assisted by AI do not perform better than radiologists without assistance 5. AI hallucinations (i.e., fabrication of information) are a major problem when humans interact with LLMs 6. LLMs can produce flawed outputs that clinicians fail to identify 7.

Of course, much of this relates to how clinicians might use AI. But what about the general public? If one in six adults are consulting AI chatbots for health information, are they getting more accurate info than basic web searches or of what a clinician might provide? That’s what a group of researchers from Oxford intended to find out. Let’s take a look at their study.

Hypothesis

The researchers did not provide a formal hypothesis.

frustrated woman on a computer

What Did They Test and How?

Study Design

The researchers conducted a study of 1,298 participants from the United Kingdom (UK). Each subject was tasked with identifying potential health conditions and a recommended disposition (course of action) in response to one of ten different medical scenarios. The scenarios were developed by a group of three doctors who unanimously agreed on the correct dispositions for each. The scenarios were then given to a distinct group of four doctors to provide differential diagnoses (Figure 1). Participants were then randomly assigned to four experimental arms. Participants in three treatment groups were provided with an LLM (GPT-4o, Llama 3, Command R+) for assistance in identifying conditions and dispositions. Participants in the control group were instructed to instead use any methods they would typically employ at home.


If you would like to continue reading...

New from Biolayne

Reps: A Biolayne Research Review

Only $12.99 per month

  • Stay up to date with monthly reviews of the latest nutrition and exercise research translated into articles that are easy for anyone to understand.
  • Receive a free copy of How To Read Research, A Biolayne Guide
  • Learn the facts from simplified research

About the author

About James Krieger
James Krieger

James has a Master's degree in Nutrition and a second Master's degree in Exercise Science He has published research in prestigious scientific journals, including the American Journal of Clinical Nutrition and the Journal of Applied Physiology, and has collaborated with notable scientists in the field like Dr. Brad Schoenfeld. He’s the former science editor for...[Continue]

More From James