Drug discovery is where this technology has the strongest claim in life sciences, and also where the claims run furthest ahead of the evidence. Both things are true at once, which makes it an awkward subject to write about honestly.
Here is a practitioner's view: what is genuinely working, what is promising but unproven, and why the bottleneck was never the part everyone is automating.
What is genuinely working
Protein structure prediction. The clearest success. Predicting three-dimensional structure from sequence went from a decades-long problem to a routine computational step, and it has changed how structural biology work is planned.
Virtual screening at scale. Computationally filtering very large compound libraries to a shortlist worth synthesising. Not new in principle, substantially better in practice.
Property and ADMET prediction. Estimating absorption, metabolism, toxicity and physicochemical properties early enough to kill a bad series before it consumes two years. Imperfect, and still useful — a prediction that is right most of the time is valuable when the alternative is synthesising everything.
Generative chemistry. Proposing novel structures against a target profile. Real, and the proposals still need a chemist to judge synthesisability and novelty.
Literature and data triage. Unglamorous and probably the highest return per pound spent — pulling signal out of internal data and published work that no team has time to read.
What is promising but not proven
Target identification. Proposing which biological target to pursue is the highest-value prediction in the industry and the hardest to validate, because the feedback loop is a decade long. Plenty of plausible candidates; little settled evidence about hit rates.
Clinical success prediction. Forecasting which molecules will survive trials. The industry would pay almost anything for this. Nobody has demonstrated it convincingly, and the base rates make it a genuinely hard statistical problem.
End-to-end discovery claims. Several companies have moved AI-originated molecules into clinical trials, which is a real milestone. What has not yet been established is whether such molecules succeed at a materially better rate than conventionally discovered ones — that evidence takes years to accumulate, and it is the only number that would settle the argument.
The honest position in late 2026 is that discovery timelines have genuinely compressed in the early stages, and the late stages — where most of the cost and nearly all of the attrition sits — look much as they did.
The bottleneck nobody is automating
Roughly nine in ten candidates that enter clinical development fail, most often on efficacy in humans or on safety that preclinical work did not predict. That failure is a biology problem, not a chemistry or computation problem.
Designing a molecule faster does not improve the odds that the target was the right one, or that the animal model predicted human response. Compressing a two-year discovery phase inside a twelve-year programme is real and worth having; it is not the order-of-magnitude change the headlines imply.
The place where better prediction would change the economics most is not generating candidates but killing bad ones earlier and more confidently — which is precisely the part that depends on evidence quality rather than model capability.
Where the regulated part begins
Most of discovery sits outside GxP, which is why teams move quickly there and are right to. The constraint arrives at the boundary: the moment a model's output supports a regulatory decision — a submission, a specification, a safety conclusion — it falls within scope and needs the evidence to match.
FDA's draft guidance on AI supporting regulatory decisions is built around exactly that distinction: the credibility you must demonstrate depends on the context of use and how much the model influences the decision. A tool that prioritises compounds for your chemists needs far less than one producing data in a filing.
Teams that get this wrong usually get it wrong late — a model used informally for two years turns out to underpin something in a submission. Deciding the boundary early is cheaper. See GxP AI validation and what regulated teams can deploy.
On the terminology: the US federal switch to "SI" changes nothing described above — see what the rename means for life sciences.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
