Trending Topics

From theory to practice: how AI is reshaping drug discovery
As part of our new “The future of life sciences R&D” series, exploring how emerging technologies – from AI and digital trials to synthetic biology and advanced data platforms – are reshaping the future of drug discovery, clinical research, and scientific innovation, we invited Cameron Ross, SVP of Generative AI at Elsevier, to share his views.
Here, Cameron explores how artificial intelligence is moving from theory to practical application across drug discovery. Drawing on developments in scientific AI and research workflows, Cameron examines where AI is already accelerating data synthesis, target identification and lead optimisation, while arguing that trusted data, explainable models and robust technical foundations are essential if organisations are to realise AI’s full potential in pharmaceutical R&D.
It takes 10 to 15 years and up to $2.8 billion to bring a drug to market. Around 80 to 90% of candidates fail in the clinic, and the early discovery stage alone accounts for 42% of total development costs. These numbers reflect a process operating under significant scientific and economic pressure.
AI offers a meaningful way to address that pressure. For many R&D and technology leaders, the question has already shifted from whether AI can help to where to focus and what foundations are needed to see results.
In drug discovery, I see three areas where AI shows significant and practical promise: data synthesis, target identification and lead optimisation. In each case, AI is not replacing scientific judgement. It is augmenting researchers, helping them cover more ground, apply criteria more consistently, and make decisions grounded in a broader evidence base. But realising these benefits at scale depends on getting the right foundations in place.
1 – Data synthesis
Every drug discovery project begins with evidence gathering: published literature, compound chemistry, disease biology, clinical trial results and proprietary data. This is a significant manual undertaking. Researchers spend an estimated 80% of their time collecting, cleaning and organising data. But even then, important connections buried across thousands of papers can still be missed.
Domain-specific large language models can search, summarise and synthesise literature at a scale that is not feasible manually, while machine learning algorithms are better suited to structured, numerical and experimental data. AI adoption for data synthesis is already accelerating: 61% of researchers now use AI to summarise research and 51% for literature reviews. But in scientific research, a summary that cannot be verified is of limited value. AI outputs must be grounded in trusted data and link back to evidence, keeping researchers in control of interpretation and validation.
2 – Target identification
Once evidence is consolidated, the next challenge is target identification: understanding the mechanism driving a disease, identifying molecular targets to affect that mechanism and prioritising candidates worth validating.
This is where AI’s pattern recognition capability is valuable. Relevant biological signals do not always sit within a single disease area. For example, asthma and some cancers may appear unrelated, but share overlapping inflammatory pathways. By integrating external databases with proprietary internal data, AI can surface these connections and rank options based on criteria such as disease association, existing literature and safety considerations.
One example is RARE Hope, a non-profit dedicated to accelerating treatments for rare neurological diseases. The team used AI-powered analysis to examine nearly 2,000 existing drugs, drawing on curated knowledge graphs and biological relationship data to score and prioritise candidates based on their interaction with disease-related targets. This helped narrow a long list of possibilities into a shortlist for expert review and validation.
3 – Lead optimisation
Lead optimisation is where thousands of potential compounds are narrowed to a workable shortlist. It is a computationally intensive challenge: improving one property can compromise another, and each design-make-test-analyse cycle is costly to run.
Machine learning models can support predictions around structure-activity relationships, ADME properties and toxicity profiles before options are tested in the lab. Generative approaches can suggest molecular designs within defined constraints. Together, these AI-assisted workflows can accelerate lead optimisation by up to 30 months.
That acceleration doesn’t remove the need for experimental work. It helps researchers focus effort where it is most likely to be productive, supporting faster learning and more efficient use of scientific capacity.
Making AI work
The potential across data synthesis, target identification and lead optimisation is clear. But potential and practice are not the same thing. Today, only 22% of researchers describe AI as trustworthy, and nearly half feel undertrained in how to use it. Closing that gap requires more than better tools. It requires organisations to build the conditions in which AI can be used reliably and with confidence. Low trust stems from inconsistent data quality, limited explainability, gaps in training and institutional inertia. There are three requirements that matter most.
Trusted data
Most scientific data was created for human interpretation, not machine reasoning. To be useful at scale, it needs to adhere to FAIR principles — Findable, Accessible, Interoperable and Reusable — so that AI systems can retrieve and connect information consistently. This also requires semantic enrichment: the ability to recognise that “GLP-1” and “glucagon-like peptide-1” refer to the same concept, for example, achieved through ontologies and controlled vocabularies. Knowledge graphs extend this further, preserving meaning across diseases, pathways, proteins and compounds, creating a foundation where data is searchable, linked and reusable.
The right architecture for the right task
Different data types require different approaches. Large language models are well-suited to literature analysis; machine learning models are often better for structural, numerical or experimental data. Retrieval-augmented generation (RAG) frameworks ground outputs in source material rather than generating unsupported conclusions. Agentic workflows will also become increasingly important. Tasks that previously ran sequentially, such as literature review, dataset search and target ranking, can be orchestrated end to end. The design-make-test-analyse cycle is the most immediate candidate for this kind of agentic coordination.
Explainability and provenance
In drug discovery, a weak target decision has consequences for timelines, investment and, ultimately, for patients waiting on treatments that do not progress. That makes traceability non-negotiable. Researchers need tools that cite their sources, are trained on peer-reviewed literature and are built to protect unpublished and commercially sensitive data, both of which drug discovery scientists routinely work with. Explainability matters here too: transparent workflows, documented reasoning and human oversight at every stage. AI should function as a co-pilot, not an autopilot, augmenting researchers’ ability to find the data they need and make well-informed decisions.
Getting the foundations right
AI will not make drug discovery automatic or remove the uncertainty inherent in scientific exploration. But it can help teams make better use of evidence, surface connections earlier and focus experimental effort more effectively across data synthesis, target identification and lead optimisation.
For R&D and technology leaders, the priority is not choosing a model. It is asking whether the foundations beneath it are ready: connected data, traceable outputs, architectures matched to the task, secure handling of proprietary and unpublished findings, explainable workflows and researchers trained to use AI critically and with confidence.
Get those foundations right and compressing that 10 to 15-year timeline stops being an aspiration. It becomes an achievable goal, one that matters not only for R&D efficiency, but for the patients waiting on the other side of that process.
