Trending Topics

The lab is digital. The data still isn’t connected.
As part of our new “The future of life sciences R&D” series, exploring how emerging technologies – from AI and digital trials to synthetic biology and advanced data platforms – are reshaping the future of drug discovery, clinical research, and scientific innovation, we invited Sean Blake, Chief Information Officer, Sapio Sciences, to share his views.
Here, Sean argues that the next phase of AI-driven research depends less on new algorithms and more on creating connected, interoperable data environments. Drawing on more than two decades of IT leadership experience in biopharma R&D, Blake examines why fragmented laboratory systems continue to limit scientific productivity and how unified data architectures could pave the way for both scientist-led and future agent-led research workflows.
According to industry non-profit The Pistoia Alliance, 80% of life sciences organizations now have cloud data platforms, 81% are using Electronic Lab Notebooks (ELNs) in place of paper notebooks, and 73% use Laboratory Information Management Systems (LIMS) to manage lab workflows and operations. These technologies are helping labs handle vast volumes of experimental, clinical, and operational data. The investment has been real and sustained.
So why are nearly two-thirds of scientists (65%) still reporting that they have had to repeat costly experiments because the data was too difficult to find?
Digitization without connection
The life sciences lab has accumulated digital systems gradually over two decades. Electronic lab notebooks replaced paper. LIMS took over sample tracking and inventory. Analytical platforms were added to handle specific data types. Each system solved a real problem at the time it was deployed. None of them were necessarily designed to share meaning with the others. The result is an environment that looks digitized but functions as a set of disconnected repositories. Data is difficult to locate, contextualize and reuse. Scientists spend time moving information between systems rather than acting on it, with more than half of scientists saying they lose significant time to exactly that kind of manual data movement, and only 5% saying they can analyze experimental results independently within their organization’s official platforms.
This fragmentation causes scientists to look for workarounds. Forty-five percent of scientists admit to using public generative AI tools through accounts they created themselves, outside any governed channel. Rather than viewing this as a rogue behavior problem, it should be a signal that the official infrastructure is not meeting demand. Scientists are finding their own solutions because the alternative is slower science.
Fragmentation that drives this is not a failure of individual systems. Each platform, considered in isolation, typically does its job. The failure is structural: systems that were never designed to share a common data model cannot share meaning, regardless of how well each one is maintained.
A data quality problem or an architecture problem?
When technology initiatives stall, the instinctive response of IT and data leaders is to improve data quality: run cleaning routines, fix formatting inconsistencies, and fill gaps in records. That work has value, but it does not address the underlying infrastructure problem.
Data cleaning corrects errors within a single system, but it does not establish shared meaning or connection across lab systems. That connectivity is required for data and related technologies to be genuinely useful at scale across research workflows.
Four patterns appear consistently in organizations where the distinction between a data quality issue and an infrastructure issue has not been addressed. The first is semantic drift: as research programs evolve and teams turn over, shared vocabularies diverge across platforms. A classification agreed on at the start of a study may carry several variant representations by the time it reaches another system. Each record is internally consistent within teams and their systems, but an inconsistency surfaces when you try to reason or interpret across different teams.
The second is the integration tax, which refers to every time a new initiative has to begin with weeks of manual data mapping and format reconciliation before any analytical work can start. That effort does not create a shared infrastructure that delivers long-term benefits. Instead, the integration effort recurs with every subsequent project, absorbing resources that should be going into the science itself.
The third is meaning lost in handoffs. When data moves between systems or research phases, contextual details that exist in the source system are silently dropped when they have no equivalent field in the destination. This means provenance chains break and outputs cannot be traced back to the conditions under which they were generated. Those losses are invisible inside any single system and only become apparent when results cannot be reproduced or audited.
The fourth is the risk of integrity erosion through duplication and divergence. When systems are disconnected, data does not remain static — it is copied, manually re-entered, and independently maintained across platforms as it moves between research phases or teams. Each copy may then evolve separately, with no reliable mechanism to detect or reconcile divergence if it occurs. A compound, sample, or subject record that begins as a single authoritative entry could accumulate conflicting versions across systems, and there may be no way to determine which is correct. That risk does not present as a formatting error that cleaning routines can identify. It may only become visible when results are compared across systems, when an audit requires a complete provenance chain, or when a finding cannot be reproduced because the conditions recorded in one system do not match those recorded in another. In regulated environments, disconnected architecture makes that assurance structurally difficult to provide, regardless of how carefully each individual system is maintained.
None of these patterns are fixed by better data hygiene. They are fixed by a different kind of infrastructure decision.
What built-in AI actually means
There is now strong appetite across the scientific community for AI as a solution to these challenges. Research found that 96 percent of scientists say instant AI-driven data analysis and visualization would be useful. But AI applied to a fragmented data environment does not resolve the fragmentation, it inherits it and makes the consequences visible faster.
The distinction that matters is between AI bolted onto existing infrastructure and AI built into a platform where the underlying data environment is unified by design. The first approach adds capability on top of the same disconnected systems. The second changes what those systems can do together.
When AI operates across a platform where LIMS, ELN, and scientific data management share a single common data model, it can act as a genuine connective layer rather than another interface scientists have to manage. Through natural language prompts, a scientist can guide each step of an experimental process while the platform calls the appropriate validated tool, retrieves the relevant information, and returns the output into a central governed record. A researcher looking for new candidate compounds, for example, could ask AI to search for structurally similar molecules using an approved modeling tool, then assess synthesizability and explore possible synthesis routes using an approved chemistry platform, all from one interface, without manually moving between systems. Each step is logged. Each output is traceable. The scientist stays in control of the decisions that matter.
This works because the data underneath it is structured to support reasoning across the full environment. The AI is not compensating for fragmentation. It is operating on a foundation designed to eliminate it.
That does not mean the path from here to there is straightforward. Moving from fragmented legacy systems toward a unified data environment involves real cost, data migration work, and significant organizational change. For most enterprises, that is a multi-year commitment, not a switch. But the architectural direction can be decided now, independently of the migration timeline. Organizations that defer the architectural question while they manage existing systems are not buying time. They are compounding the problem and narrowing the options available when AI deployment becomes unavoidable.
From scientist-led to agent-led
Today, the practical model is scientist-led: the researcher decides what to ask next, approves each output, and directs the sequence of steps. The AI coordinates; the scientist governs.
Longer term, as confidence in AI systems grows and models demonstrate sustained reliability, a shift toward more autonomous agent-led operations becomes credible. In that model, the scientist defines a hypothesis and the system orchestrates the tools, data sources, and analytical steps needed to test it, returning evidence or flagging decisions at the points where human review matters most. Teams can explore multiple research routes virtually before any physical lab work begins, reducing the cost of failed experiments and directing effort toward the approaches most likely to succeed.
That model requires a clear foundation. Agent-led operations do not bypass the need for a solid data foundation. They depend on it more heavily than scientist-directed work does, because there is no human in the loop at each step to catch errors or reinterpret ambiguous outputs. Consistent identifiers, preserved provenance chains, and shared semantic structure are not prerequisites that become less important as AI becomes more capable. They become more important. Autonomous operations running on fragmented data do not produce slower errors. They produce faster ones at a greater scale with less visibility.
The scientist still makes the critical decisions in either model. Their judgment is reserved for the moments where it counts: validating outputs, challenging assumptions, and deciding whether the evidence is strong enough to move forward. What changes is the scope of what the system can prepare and coordinate on their behalf.
A joint decision, not a sequenced one
The risk in framing this as a data problem is that it becomes justification for deferring AI investment: fix the infrastructure first, then revisit AI when the foundation is ready. That sequence is the wrong one, and it tends to mean neither problem gets solved.
The more productive framing is a joint architectural commitment made by IT and scientific leadership together. The infrastructure choices that determine what AI can do are as much scientific and operational decisions as technical ones. What data needs to mean, how it should be governed, and which systems need to share a common foundation: these are questions that require both perspectives to answer well and that neither team can resolve alone.
Organizations that make that decision jointly and make it early build infrastructure that compounds in value. Those that treat it as an IT sequencing problem tend to find themselves renegotiating the same architectural questions with every new AI initiative that stalls.
The lab is already digital. Making it genuinely connected is the work that determines what comes next.
