AI in Drug Discovery: Moving Beyond Molecule Identification to Clinical Success
Explore how AI in drug discovery is moving beyond molecule identification to improve target discovery, drug development, clinical trials, and patient outcomes.
Blog
AI in Drug Discovery: Moving Beyond Molecule Identification to Clinical Success
AI in Drug Discovery: From Molecules to Clinical Success
Explore how AI in drug discovery is moving beyond molecule identification to improve target discovery, drug development, clinical trials, and patient outcomes.
AI in Drug Discovery: Moving Beyond Molecule Identification to Clinical Success
Published: Sept 26th, 2026
Only around 14.3 percent of investigational drugs that enter Phase 1 trials go on to reach regulatory registration, and the figure is markedly lower still in oncology and central nervous system indications (Schuhmacher et al., Drug Discovery Today, 2025). The attrition is concentrated well after the molecule has been identified. A compound can clear target validation, perform predictably in preclinical models, and still fail in the clinic because the biomarker used to track response did not generalise, the trial could not recruit an eligible population fast enough, or a safety signal surfaced too late to adjust the protocol.
Molecule identification has become the tractable part of the problem. Machine learning models can now screen chemical space, predict binding affinity, and rank candidate compounds at a pace no manual medicinal chemistry team could match. That capability is well established and reasonably well served by existing tooling. What remains underdeveloped is the application of AI in drug discovery and development across everything that happens after a candidate leaves the lab: biomarker validation, patient stratification, trial feasibility modelling, and the regulatory scaffolding required to make any of that defensible.
Where AI in drug discovery currently performs well
Target identification and validation
Target identification is the stage most commonly associated with AI in drug discovery, and for a reason that's easy to see once you look at the inputs: they're already structured. Genomic, proteomic, and literature data lend themselves to models that can surface candidate disease targets faster than a manual review, ranking them against known pathway interactions and prior evidence. But a target that scores well computationally isn't automatically a target that's been validated. Statistical association within a dataset is not biological proof, and plenty of promising-looking targets fall apart once tested against real physiology.
Biomarker discovery
Biomarkers are arguably where the use of AI in drug discovery has the most direct bearing on whether a program actually succeeds, as opposed to just moving faster in the early stages. They decide who's eligible for a trial, how efficacy gets measured, and eventually what shows up on a prescribing label. Get one wrong, and that error doesn't stay contained, it propagates through every stage that follows.
This is part of why time-series approaches, including Neural ordinary differential equation (Neural ODE) architectures, have gained traction on longitudinal clinical data. Disease progression and treatment response don't unfold at neat, fixed checkpoints, they evolve continuously, and models built to handle continuous-time dynamics can pick up on biomarker candidates that a snapshot-based statistical method would simply miss. Crucially, this can happen while a trial is still recruiting, not months later during retrospective analysis.
Patient stratification and trial feasibility
Once development moves into the clinical phase, this is a central answer to how is AI used in drug discovery. Genomics, electronic health records, imaging, and claims data can be pulled together to model which patient subgroups are likely to respond to a given mechanism, and whether a given region even has enough eligible patients to hit enrolment timelines. A surprising amount of trial delay traces back to inclusion criteria nobody stress-tested against real-world population availability before the protocol was locked in.
Literature and real-world evidence synthesis
There's a lot of relevant information sitting in published literature, competitor trial registrations, and real-world evidence datasets that simply never gets reviewed, not because it isn't useful, but because there's too much of it for a team to process by hand. Natural language processing built for scientific and clinical text changes that math: it can surface prior art, competing mechanisms, and early safety signals that would otherwise take months of manual digging to compile.
Deployment and lifecycle management
Here's a distinction that gets glossed over too often: a model that performs well in a research environment and a model that runs reliably inside a pharmaceutical company's production infrastructure are not the same thing, and they're governed by different requirements. Getting from one to the other takes real MLOps discipline, version control, performance monitoring, defined retraining triggers, built around the data governance constraints and regulatory obligations that already exist, not bolted onto a generic machine learning deployment stack.
Why faster discovery has not yet meant more approved drugs
Does AI in drug discovery actually improve clinical success, or just discovery-stage speed? The honest answer is that the evidence cuts both ways, and it's worth laying out plainly rather than rounding to either optimism or scepticism. Nature Biotechnology published an analysis showing AI-discovered candidates advancing through Phase 1 at an 80 to 90 percent success rate, well above the industry's historical 40 to 65 percent range (Drug Discovery News, 2026, summarising Jayatunga et al.). That's a real, measurable signal, not a marketing claim. But it hasn't carried through to later trials. A follow-up analysis in NPJ Drug Discovery found that while AI-derived molecules do trend toward better Phase 1 safety and tolerability, their Phase 2 proof-of-concept success rates land right around the historical norm of roughly 40 percent. And a 2025 review in Clinical Pharmacology and Therapeutics goes further, noting that at the time of that review, no novel AI-discovered drug had reached full regulatory approval.
Why the gap? Probably because AI's current strengths, generating and prioritising hypotheses, predicting molecular structure and properties, triaging candidates on safety, line up with the parts of discovery that were already reasonably well understood. The parts that fail most expensively, whether a target genuinely drives human disease, whether a biomarker actually generalises to the treated population, are exactly the areas this article keeps circling back to as underdeveloped relative to molecule-identification tooling. A 2026 review in Genes & Diseases lands on something similar from the development side: AI now touches every stage from early lead screening through clinical trial implementation, but data reliability, verification of model predictions, and regulatory supervision are still what's holding wider adoption back.
The regulatory landscape is still forming, and its boundaries matter
Three developments define where things currently stand on AI in drug development, and each one has real implications for how a program should be designed.
Start with the fact that the FDA's discussion paper came before any formal guidance. In May 2023, and again in a February 2025 revision, the FDA's Center for Drug Evaluation and Research published "Using Artificial Intelligence and Machine Learning in the Development of Drug and Biological Products." It was meant to solicit stakeholder feedback, not establish binding policy, and it's worth remembering that distinction when people cite it as if it were a rule.
The actual first formal draft guidance came in January 2025. "Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products" [FDA-2024-D-4689] lays out a risk-based credibility assessment framework built around a defined context of use (COU), with seven steps for planning, gathering, and documenting evidence of model credibility. A 2026 Genes & Diseases review summarises this independently as requiring sponsors to specify a clear context of use and maintain human oversight and prospective monitoring, which lines up with how the FDA itself describes it. Notably, this guidance explicitly excludes AI applications in drug discovery and general operational efficiency from its scope, since it applies only where an AI model generates information intended to support a regulatory decision on safety, effectiveness, or quality. That scoping choice leaves discovery-stage and biomarker-validation work in a comparatively undefined regulatory position, even though decisions made at that stage shape everything that follows.
The EMA takes a wider view. Its "Reflection Paper on the Use of Artificial Intelligence in the Medicinal Product Lifecycle," adopted jointly by the CHMP and CVMP on 30 September 2024, covers AI and machine learning across the entire product lifecycle, from discovery all the way through post-authorisation pharmacovigilance. It recommends developers seek early regulatory engagement, such as EMA qualification of novel methodologies, whenever an AI or ML system is expected to influence a product's benefit-risk balance. The same 2026 Genes & Diseases review points out that the EMA paper stresses continuous monitoring for model drift, since the population a model was trained on can shift after approval. If you're operating across both US and EU markets, this divergence at the discovery stage is something to plan around, not something to discover later.
None of this replaces existing validation obligations, either. Where an AI model supports a GxP-regulated activity, GAMP 5 Second Edition (2022), which added specific guidance on machine learning systems, and 21 CFR Part 11 audit trail and record integrity controls still apply. A model that keeps updating itself raises a change control question that a static, locked model never had to answer, and that's something to work out during validation planning, not after the model's already live in a regulated workflow.
Why the scope beyond molecule identification matters
A faster candidate-generation pipeline feeding into an otherwise unchanged, failure-prone clinical process doesn't move the number that actually matters, time to a safe, effective, approved therapy. The programmes seeing real returns from AI treat it as one continuous thread running from target validation through biomarker development, trial design, and regulatory submission, rather than something applied at the discovery stage and then set aside.
That's also where the real implementation difficulty sits. Building a decent discovery-stage model is, at this point, a fairly well-solved problem. Validating that model's outputs against real clinical data, wiring it into existing pharmaceutical data infrastructure, and actually earning the trust of a clinical team making prescribing or protocol decisions, that's a much harder, far less standardised job. Life sciences work that specifically targets this gap, using Neural ODE modelling on longitudinal clinical data for biomarker discovery, combining literature and multimodal data mining for target and biomarker validation, building MLOps infrastructure that lets pharmaceutical teams run several production models without tearing out their existing systems, gives a sense of where the discipline is actually heading. Reveal HealthTech's life sciences practice is one example of a team working specifically in this space, with case work spanning biomarker modelling on time-series clinical data and MLOps deployment for pharmaceutical clients running multiple production models at once.
Frequently Asked Questions
AI in drug discovery refers to the application of machine learning and related computational methods to identify and prioritise candidate compounds, most commonly through analysis of chemical, genomic, and proteomic data to predict binding affinity, target engagement, and early safety liabilities.
Beyond identifying and validating targets, AI is used for biomarker discovery from longitudinal clinical data, patient stratification and trial feasibility modelling using multimodal datasets, literature and real-world evidence synthesis, and the MLOps infrastructure required to run validated models reliably in production.
Drug discovery refers to the earlier, preclinical stage of identifying and validating candidate molecules and targets. Drug development encompasses the subsequent stages, including clinical trial design, patient recruitment, biomarker validation, and regulatory submission, where AI in drug discovery and development increasingly operates as a continuous process rather than two separate applications.
The FDA's January 2025 draft guidance on AI in regulatory decision-making explicitly excludes AI applications used in drug discovery and general operational efficiency from its scope, applying instead to AI models that generate information supporting a regulatory decision on safety, effectiveness, or quality later in development.
No. The EMA's September 2024 reflection paper addresses AI and machine learning across the entire medicinal product lifecycle, including the discovery stage, which is a broader scope than the FDA's current draft guidance.
Where a model supports a GxP-regulated activity, existing computerised system validation requirements apply, including GAMP 5 Second Edition guidance on machine learning systems and 21 CFR Part 11 audit trail controls. Continuously updating models require a defined change control process.
Blog
How AI is transforming medical affairs through real-time scientific intelligence
How AI Is Transforming Medical Affairs with Real Time Scientific Intelligence
Discover how AI is transforming medical affairs with real-time scientific intelligence, enabling faster insights, smarter decisions, and more effective healthcare strategies.
AI in Drug Discovery: Moving Beyond Molecule Identification to Clinical Success
AI in Drug Discovery: From Molecules to Clinical Success
Explore how AI in drug discovery is moving beyond molecule identification to improve target discovery, drug development, clinical trials, and patient outcomes.
As AI scales in health systems, models and data need a shared operating layer.
As healthcare AI scales, healthcare organizations need more than high-performing models. They need an operating layer that integrates data, governance, workflows and human oversight into one coherent, trustworthy environment.