AI drug discovery hits a data wall: models trained on success can't learn from failure
AI is identifying drug candidates faster than labs can test them, exposing a bottleneck few discuss: models trained on published successes cannot learn from the failures companies keep private, raising questions about the reliability of AI-generated therapeutics and the integrity of the data feeding them.

AI Drug Discovery's Survivorship Problem: Models Built on Wins Can't Learn from Losses
AI drug discovery has a data problem that no amount of compute can fix. The models generating promising drug candidates at unprecedented speed are trained almost entirely on published successes. The failures, the dead ends, the compounds that looked promising and flopped, sit in proprietary lab notebooks no one sees. This creates a survivorship bias baked so deeply into the training data that the industry cannot measure how much it distorts predictions.
A machine learning model can produce a binding score for a novel molecule with three decimal places of apparent precision. That number is only as reliable as the data behind it. When the training set contains no examples of similar molecules that failed for unanticipated reasons, the model has no mechanism to flag the risk. It projects confidence its training data never earned.
Paul Belcher, director of protein research strategy at Cytiva, a global life sciences company, describes this as hitting a "data wall." Because models draw on the same publicly available datasets, they converge on similar answers with diminishing returns. The datasets were never built for machine learning in the first place, lacking the structure, labeling, and diversity needed to keep predictions accurate and unbiased 1.
The deeper problem is publication bias, a structural feature of scientific research documented for decades. Papers with statistically significant results are roughly three times more likely to be published than those with null findings. Negative results get filed away in what psychologist Robert Rosenthal in 1979 called the "file drawer problem" 2. In pharmaceutical research, competitive secrecy compounds the effect. A failed compound is proprietary intelligence. Publishing it would hand competitors a roadmap of what to avoid.
Belcher puts it plainly: nobody shares their failures. The data that would most improve model reliability remains locked in lab notebooks, never used to inform future research 1. He jokes that there should be a "journal of negative data." Without failure data, a model that only knows what succeeds cannot distinguish a genuine hit from a statistical mirage.
The economics sharpen the stakes. The inflation-adjusted cost of developing a new drug has roughly doubled every nine years since the 1950s, a trend Jack Scannell and colleagues dubbed Eroom's Law (Moore's Law backwards) in a 2012 paper in Nature Reviews Drug Discovery 3. Bringing a drug to market takes 10 to 15 years and costs $1 billion to $2.5 billion, with failure rates above 90 percent
1. AI is the industry's primary bet on reversing that trajectory by identifying candidates faster and eliminating losers earlier.
But AI can only eliminate losers it has learned to recognize. If the training data systematically excludes failures, the model's blind spots are not random. They are predictable in exactly the way that is most dangerous. A model will be most confident precisely where it has the least information, because it has never encountered a counterexample.
The data problem runs in both directions. On one side, failure data is absent. On the other, the success data that does exist may be fabricated. Research by Elisabeth Bik, a Dutch microbiologist and scientific integrity consultant who received the 2021 John Maddox Prize for exposing threats to research integrity, found that approximately 4 percent of biomedical papers contain duplicated or manipulated images 1. Her separate analysis of 960 papers in Molecular and Cellular Biology found 6.1 percent contained inappropriately duplicated images
4. That work predated generative AI tools that make fabrication trivial. The models sit on a dataset that fails in two directions at once: it is missing every failure, and some of what remains has been manipulated. Contaminated positives plus absent negatives means models receive bad information from both ends of the distribution.
The MIT Technology Review piece, produced in partnership with Cytiva, frames the solution as lab automation: integrated, autonomous systems that cycle through prediction, testing, and optimization, feeding results back into models to close the loop 1. Belcher describes these as "labs-in-the-loop" operating around the clock.
Lab automation solves the velocity problem. It does not solve the data-sharing problem. The failure data needed to retrain models sits with rival companies that have every commercial reason to keep it private. No single vendor's integrated system fixes that. It requires collective action: pre-competitive data-sharing agreements, regulatory mandates for reporting negative results, or industry consortia that treat failure data as shared infrastructure rather than competitive intelligence.
The bottleneck in AI drug discovery is not finding molecules. AI can already generate candidates faster than labs can test them. The bottleneck is knowing which candidates to trust. Trust depends on data the industry has structural reasons not to share. Every AI-generated drug candidate carries a risk premium that current models cannot calculate, because the information needed to price it sits in someone else's vault.
Until failure data becomes shareable, the ceiling on AI drug discovery is not computational. It is institutional.
References
Cite this story
ProvenBrief (2026). "AI drug discovery hits a data wall: models trained on success can't learn from failure." ProvenBrief. https://provenbrief.com/story/ai-drug-discovery-hits-a-data-wall-models-trained-on-success-can-t-learn-from-fa
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.