🌿freegardner

Synapse

AI chatbots miss most relevant medical studies

19 Aug 2026 · via Rss.arxiv

AI chatbots miss most relevant medical studies

AI chatbots miss most relevant medical studies

The Cochrane Library is the gold standard of medical evidence. For decades, its systematic reviews have been assembled by trained experts who spend months combing through thousands of studies, applying strict inclusion criteria, and cross-checking every citation. That human labor is slow, expensive, and increasingly seen as a bottleneck. So when researchers at a major university recently asked three AI chatbots to perform the same literature retrieval task, they weren’t just testing accuracy. They were testing whether the entire profession of evidence synthesis still has a job.

The results, published as an arXiv preprint in 2026, are unsettling in a way that has nothing to do with robots taking over. The chatbots retrieved an average of 39.2 percent of the studies that Cochrane’s human experts had deemed relevant, with a standard deviation of 29.8 percent That number alone tells a story of mediocrity. But the deeper finding is more insidious: the AI systems didn’t just retrieve studies poorly. They retrieved them differently, and that difference reveals a bias that no amount of fine-tuning will fix.

The Size Bias Nobody Programmed

When the researchers dug into which studies the chatbots chose to cite, one factor stood out above all others: sample size. A trial with 10,000 participants was nearly twice as likely to be retrieved as a trial with 1,000, even when the smaller study was methodologically superior. This wasn’t a design choice. No prompt instructed the models to favor large trials. It emerged organically from the training data, where big studies dominate headlines, abstracts, and citation networks.

This is the quiet danger of AI replacing human judgment. A human librarian might be biased by recency or by the prestige of a journal, but they can be trained to correct for those tendencies. An AI system carries its biases invisibly, baked into billions of parameters that no one fully understands. The chatbots weren’t deliberately ignoring small but rigorous trials. They simply had learned, from the statistical texture of the literature itself, that size matters more than quality.

The consequences are not abstract. In clinical medicine, small trials often address rare conditions or neglected populations. A chatbot that systematically overlooks these studies isn’t just making an academic error. It’s making a recommendation that could shape treatment decisions for patients who are already underserved.

Who Asks Matters More Than What You Ask

The study also found that the role of the person asking the question changed the results. When the researchers prompted the chatbots as evidence-synthesis experts, retrieval rates improved significantly compared to when they asked as patients or clinicians. This might sound like good news—the AI performs better when it knows it’s talking to a professional. But flip that finding around and it becomes deeply troubling.

The person most likely to benefit from AI-assisted literature search is the patient who doesn’t know the technical vocabulary, who can’t articulate their question in the language of systematic review methodology, and who doesn’t realize that the quality of the answer depends on how the question is framed. The chatbot gives its worst performance to exactly the people who need it most. This is automation deepening inequality rather than reducing it.

The gap between roles—42.8 percent recall for researchers versus 36.1 percent for patients—might seem modest. But in a field where a single missed study can change a clinical recommendation, that six-point difference represents real harm. The patient who asks “what should I take for my arthritis?” gets a worse answer than the researcher who asks “what does the evidence say about the efficacy of NSAIDs compared to acetaminophen for osteoarthritis?” The AI isn’t just retrieving studies. It’s silently deciding who deserves a better answer.

AI chatbots miss most relevant medical studies (Bild 1)

The False Comfort of Citation Counts

One of the most seductive features of modern chatbots is their ability to provide citations. Unlike earlier AI systems that hallucinated sources, current models like Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5 can point to real papers. The study confirmed this: only 5.0 percent of cited studies were ones that Cochrane had explicitly excluded, meaning the chatbots rarely fabricated references On the surface, this looks like progress.

But a citation is not a judgment. A chatbot that correctly identifies a study as real is not the same as a chatbot that correctly identifies a study as relevant. The excluded-study rate of 5.0 percent is reassuring only if you ignore the 60.8 percent of relevant studies that were never retrieved at all. The AI is doing the easy part of the job—finding papers that exist—while failing at the hard part—knowing which papers matter.

This is the classic pattern of automation replacing judgment with retrieval. A librarian who retrieves 39.2 percent of the relevant literature would be fired. A chatbot that does the same is celebrated for not hallucinating. The bar for AI performance has been set so low that mere competence looks like excellence.

The Uneven Playing Field of Models

The differences between the three chatbots were stark. ChatGPT achieved a recall of 63.1 percent, nearly double that of Claude at 37.0 percent, and more than triple that of Gemini at 17.3 percent. These are not subtle variations. They represent fundamentally different approaches to the same task, trained on different data, optimized for different objectives.

The problem is that users don’t know which model they’re getting. A clinician who opens a chat interface in 2026 has no way of knowing whether they’re using a system that will find 63 percent of the relevant studies or one that will find 17 percent. The interface looks the same. The confidence of the language is identical. The only difference is invisible to the user.

This is the paradox of AI replacing human expertise. When a librarian does a poor job, you can see it. The search takes too long, the results are incomplete, the librarian looks uncertain. When a chatbot does a poor job, it sounds exactly as confident as when it does a good job. The technology has eliminated the external signals that once allowed us to calibrate our trust.

The Logic of the Larger Trial

Why do the chatbots prefer larger studies? The statistical analysis in the paper offers a clue: sample size was the only independent predictor of retrieval, controlling for publication year, citation count, and open-access status. This suggests the bias is not about visibility or prestige. It’s about something more fundamental in how language models learn from text.

Large trials are discussed differently. They generate more news coverage, more commentary, more follow-up studies. Their results are summarized in more places, with more redundancy. A language model trained on this corpus naturally assigns them higher probability. The chatbot isn’t making a rational choice about study quality. It’s reproducing the statistical regularities of the literature itself.

AI chatbots miss most relevant medical studies (Bild 2)

This creates a feedback loop. As AI systems increasingly mediate access to medical knowledge, they will drive more attention to large trials. Those trials will get more citations, more coverage, more discussion. The next generation of AI will be trained on this distorted corpus and will amplify the bias further. Small trials, already marginalized, will become nearly invisible.

The Human Cost of Efficiency

There’s a temptation to see this as a temporary problem, a bug that will be fixed with better prompts or more training data. But the study suggests something more structural. The researchers varied the user role, repeated the queries four times, and used 20 different clinical questions. The bias persisted across all conditions. This isn’t a glitch. It’s a feature of how these systems work.

The deeper issue is that AI doesn’t just replace the librarian—it replaces the entire ecosystem of judgment that surrounds the librarian. A human expert brings context, experience, and the ability to say “this study is small but important.” A chatbot brings statistical probability. In a world where speed and efficiency are valued above all, the chatbot wins. But the cost is a systematic erosion of the very qualities that make medical evidence trustworthy.

The image that stays with me is not the 39.2 percent recall rate or the sample-size odds ratio of 1.80. It’s the patient, sitting alone with a chatbot, asking about a rare condition, receiving a confident answer based on the largest trials available—which happen to be the least relevant to their situation. The chatbot will cite its sources. The citations will be real. And the answer will still be wrong.

That’s the future we’re building. Not a world where AI makes us smarter, but a world where the appearance of expertise replaces the substance of it. The librarians are gone. The chatbots are confident. And the small trials, with their small sample sizes and their crucial insights, are waiting for someone to ask the right question.

No one will.


Sources

1. Claude Sonnet 5

2. Gemini 3.1 Pro

← back to the garden