Why a closed-loop approach is the only way AI can deliver true value to life sciences R&D

As part of our new “The future of life sciences R&D” series, exploring how emerging technologies – from AI and digital trials to synthetic biology and advanced data platforms – are reshaping the future of drug discovery, clinical research, and scientific innovation, we invited Guy Levy-Yurista, CEO of Imperagen, to share his views.

Here, Guy argues that the real challenge facing AI in life sciences is not the models themselves, but the quality and structure of the underlying data. Drawing on his experience scaling AI, SaaS, and deep-tech businesses, Guy explores why a “closed-loop” approach – where experimental data continuously feeds computational models – is essential if AI is to deliver meaningful, real-world value in biotech and R&D.


The AI revolution in life sciences R&D is real, but true success will only come if we solve the data problem.   

Like a lot of industries, AI has promised much when it comes to life sciences R&D, yet the reality is often very different. The significant efficiency gains and potential for optimisation has yet to materialise. And the main culprit is the underlying data.

I’ve spent my career scaling deep tech businesses, and I’ve seen this pattern before when a powerful new technology arrives, everyone rushes to apply it, results disappoint, and people blame the technology. But technology itself usually is not the problem, it’s our approach. 

Many AI approaches in drug discovery, molecule design, and biological engineering depend on data that was never generated for the specific question being asked. Public databases, sparse historical datasets, inconsistent assay conditions, unlabeled failures, and sequences pulled from wherever they can be found all have value. The AI models training on this data produce outputs that look spectacular on a screen, but fall apart the minute someone takes them into the lab. In silico confidence, in vitro catastrophe.

That gap between AI prediction and reality is where hundreds of millions of dollars in R&D spend goes to die every year. 

So, what’s the big fix? 

The only way to generate AI outputs in science that reflect reality is to own the data generation layer entirely. We need to start building rich datasets – experiment by experiment – and feed it into computational models. Each experiment tells you what worked and what didn’t; both sets of this information are valuable and needed to train the AI models

This is what we call a “closed loop,” and the only way to train AI models for life sciences. 

What does it look like in reality? 

Biology is intensely contextual. For example, in enzyme engineering, it is important to understand on an atomic level how an enzyme performs on a specific substrate, in a specific reaction, under specific conditions, against a specific target. Only when you understand the chemistry of the reaction, you can build atomistic models, run quantum-mechanical calculations to understand the energy barriers, and screen millions of mutation combinations computationally to find few hundred variants to test in a lab. 

Testing those variants in the lab gives you information on what worked and what failed. It is important to capture that information and feed it back to computational models. 

Why failures are actually data

That point is easy to underestimate. In many experimental programs, failures are treated as dead ends, but they are a signal, an important one. A model trained only on winners gets a distorted view of the world. A model trained on both successful and unsuccessful variants has a much better chance of becoming useful. 

Generating the context rich data and continuously feeding it back into AI, makes AI predictions more realistic and useful in further research. The compounding effect of this approach is what makes it so powerful. Once the model has started to understand the idiosyncrasies of your specific system, you’re building an increasingly intelligent engine that gets better with every cycle. 

The shift to data-generation engines

What does this mean for life sciences R&D partnerships? It means the old model of handing a dataset to an AI vendor and waiting for a candidate list is finished. The companies that will win the next decade of biotech are the ones that treat their wet labs as data-generation engines, not validation facilities. Every experiment becomes an investment in model intelligence. Every result – positive or negative – feeds the next prediction. 

At Imperagen, we’re building what we call a Large Reaction Model for biocatalysis. Its purpose is to help engineer enzymes that perform in real world conditions.  Better enzyme engineering can transform how we manufacture medicines, chemicals, materials, and sustainable products. It can help replace inefficient or environmentally harmful processes with cleaner biological alternatives. It can make production more selective, scalable, and resilient.  

Our ambition is to create a foundational AI system for how enzymes are engineered, but the way to get there is by generating huge amounts of highly specific, high-quality data. 

The AI promise in life sciences is real,  the opportunity is enormous, but the path to value will come from closing the loop between AI prediction and reality.

About The Author

Avatar photo
Ricardo Oliveira

Ricardo Oliveira is a Senior Director at TechFinitive, where he frequently collaborates with TechFinitive's editorial team to write and produce content. He's based in Sydney, Australia.

Read more from this author.

We take journalism seriously. To learn more on why you should trust us, head to our editorial guidelines page or meet our team.