🌿freegardner

Science

AI Drug Discovery Faces Data Trust Gap

19 Aug 2026 · via Pharmexec

AI Drug Discovery Faces Data Trust Gap

AI Drug Discovery Faces Data Trust Gap

The pharmaceutical industry is heading toward a striking milestone: $25 billion spent on artificial intelligence by 2030. That figure promises faster discovery, lower attrition rates, and sharper decisions at every stage of the drug pipeline. Yet a quieter set of numbers tells a different story, one that should give every biopharma executive pause before signing off on the next round of funding.

According to a 2024 Drug Discovery News survey, AI adoption in pharma drops sharply in the very domains where discovery depends on it most. Generative design sits at 42% adoption. Biomarker analysis follows at 40%. ADME prediction — the study of how drugs are absorbed, distributed, metabolized, and excreted — lags even further at 29%. The bottleneck is not the models themselves, according to the survey. It is the fragmented data environment underneath them.

The gap between what leaders fund and what bench scientists actually use represents the hidden risk behind one of the largest technology bets the industry has ever made. If the scientists at the bench do not trust the tools, the promised return on that investment never materializes. The money flows, the models run, and the results sit unused.

This trust gap has a clear cause, and it is not the stubbornness of researchers who refuse to embrace new technology. AI systems are being built on incomplete, unvalidated, and untraceable data. Scientists cannot defend such outputs in a review, cannot reproduce them in a lab, and cannot stake a program decision on them. The next phase of AI investment must shift focus from bigger models to data readiness — aggregating, normalizing, and harmonizing internal and external sources into an authoritative, AI-ready database.

Why Scientists Turn Away From Confident Answers

Executives might be tempted to interpret slow AI adoption inside R&D as a cultural problem. The workforce needs training, the thinking goes, or persuasion to embrace new tools. Those factors may play a role, but the bigger issue is that scientists hesitate to rely on AI outputs for reasons fully aligned with their roles, responsibilities, and training.

Scientific work carries a built-in accountability standard. Every conclusion must be defensible. Every result must be reproducible. Every claim must be traceable to a source. This standard is not a preference — it is how science operates and how programs move forward without backtracking later. AI outputs do not meet that standard by default.

A confident-sounding answer that turns out to be wrong fails the professional bar scientists are trained to apply. An answer whose source cannot be located fails the same bar. A scientist who hesitates to act on such an output is not resisting technology. They are doing what their training requires: refusing to stake a program decision on something they cannot defend, reproduce, or trace.

AI Drug Discovery Faces Data Trust Gap (Bild 1)

Skepticism in this context is not the obstacle to AI adoption. It is the standard the technology must clear. The scientists are not the problem — they are the gatekeepers of quality, and their caution reflects the reality that AI has not yet earned its place at the bench.

The trust issues playing out across science and healthcare become more pronounced in drug discovery, and the reason has less to do with AI itself than with what AI is asked to learn from. Broad models trained on the open web were not created with the complexity of drug discovery data in mind. When these models are pointed at specialized scientific questions, errors are introduced.

Drug discovery depends on domain-specific data structures, controlled vocabularies, and annotation standards that general-purpose datasets do not reflect. A single chemical compound can appear under thousands of different names across databases and publications. AI models trained on that inconsistency inherit the confusion and carry it forward into every output they generate.

Outputs that look authoritative on the surface can mask a foundational mismatch between the question being asked and the data the model has to work with. The surface polish of a well-written AI answer hides the shaky ground beneath it. Scientists who dig into the sources often find that the model was working with fragments, duplicates, and mislabeled entries from the start.

The problem compounds because scientists are not trained to catch the failure modes of a language model, and they should not have to be. Their training is in chemistry, biology, and medicine — not in auditing statistical text generators. When a subtly wrong AI output enters a drug discovery workflow, it can move through a program before anyone notices.

The downstream effects are severe. Wasted time on results that cannot be reproduced. Budget spent on paths that lead nowhere. Delays when compromised work has to be repeated from scratch. A 2023 Nature Medicine commentary documented how a single AI-introduced error in a drug repurposing screen propagated through multiple validation steps, wasting months of laboratory work before being caught. The case illustrates how quickly one bad output can spread when the underlying data is not built for the work being done

This is what makes the trust gap so consequential in the pharmaceutical industry. The cost of misplaced trust in drug discovery is far higher than in almost any other domain where AI is being deployed. A wrong recommendation in a marketing campaign costs money. A wrong recommendation in drug discovery can cost years and lives.

The Data Foundation That Changes Everything

Addressing the trust gap comes down to three components. The first is a knowledge base scientists can trust. The second is an AI system that leverages content appropriately. The third is traceability that lets scientists see exactly where an answer came from before they act on it.

AI Drug Discovery Faces Data Trust Gap (Bild 2)

Trust starts with the underlying content. When AI draws from a rigorously human-curated knowledge base, the foundation can support the work scientists need it to do. Such a knowledge base aggregates scientific information across literature, patents, and structured databases. It normalizes everything under consistent vocabularies and identifiers. It harmonizes the results by scientific experts into an authoritative, AI-ready data asset.

The integrity of an output cannot exceed the integrity of the data feeding it. In drug discovery, the difference between curated and uncurated content is the difference between an answer worth acting on and one that has to be re-verified. Curated content has been checked, deduplicated, and standardized. Uncurated content is a swamp of conflicting names, overlapping entries, and unresolved ambiguities.

A trusted foundation only matters if the system built on top of it can use it well. Drug discovery questions are rarely answered in a single retrieval step. They require breaking a complex problem into targeted, multi-step search paths and applying domain-specific reasoning at each one. A general-purpose model that tries to answer in one leap will stumble.

AI designed by scientists for R&D approaches scientific questions differently. The workflow mirrors how a scientist would actually approach the problem: breaking it down, checking each step, and building toward an answer. This produces more nuanced and defensible answers than a general-purpose model because the process is structured the way science itself is structured.

The final component is traceability. None of this is useful if the scientist cannot see where a given answer originated. A recommendation a scientist cannot trace is a recommendation they cannot defend. When an output cannot be verified, scientists must manually re-verify the finding before acting on it, adding hours and days to workflows that AI was supposed to accelerate.

The verification burden erodes the return on investment for the tool entirely. A tool that requires as much manual checking as doing the work itself has no value. The promise of AI was to remove the verification burden, not to relocate it.

The conversation around AI in pharma has been dominated by questions about what models can do, how large they can get, and how much data they can be trained on. The more important question — and the one that will determine whether any of this investment pays off — is whether scientists trust the outputs enough to act on them. That trust, as the adoption numbers show, has not yet been earned.

That is not a question scientists can answer on their own. It is a question for the executives funding these initiatives, who must decide whether to invest in data infrastructure with the same seriousness they have applied to model development. Biopharma companies that want value from their AI investments must recognize that data readiness matters more than data volume. Pouring more general scientific content into a model does not make it better at drug discovery. It often makes the inconsistency problem worse.

The path forward requires a shift in attention. The next phase of AI investment should not just be about bigger models and broader data. It has to be about data readiness — building the authoritative, AI-ready foundation that scientists can trust, defend, and act on. Without those fundamentals, AI investments to drive drug discovery will continue to fall short of expectations, and the $25 billion bet will yield less than its promise.

The question now is whether the industry will learn this lesson before the next wave of spending, or after the next wave of disappointing results. The $25 billion commitment is already in motion. The scientists are waiting to see whether the tools built on that investment will finally meet the standard their work demands — and whether the data underneath them will finally be worthy of trust.

← back to the garden