Trending Topics

Your AI projects will stall without these data partnerships
This article is part of our Opinions section, where we invite industry professionals to share their views on the most pressing technology questions of our time. Here, Claus Wolthers, Director of Technology Partnerships and Ecosystems at Komprise, explores why AI projects increasingly depend on more than sophisticated models and infrastructure.

Most enterprises now treat AI as a priority. Far fewer have solved the problem that decides whether AI works: the data underneath it. AI models depend on unstructured data: the documents, images, audio, video, chat logs, and sensor output that carry the context CXOs want and information workers need. That data is the real dealbreaker for enterprise AI, and it has become a complex barrier to AI moving from pilot to production and achieving measurable ROI.
Solving the data puzzle for AI is a widespread, vexing issue for CIOs; unstructured data is enormous, messy and lacking quality, expensive to move and scattered across dozens of applications, storage systems and clouds. Less than 1% of unstructured data is in an AI-ready format, according to IBM, making it largely dark to AI.
Gartner predicts organisations will abandon 60% of AI projects that are not supported by AI-ready data, while 57% of data leaders view data reliability as a key barrier to moving AI projects from pilot to production, according to 2026 research from Informatica.
There is no silver bullet or single tool that makes the problem disappear: this is heavy, ongoing work across IT and departments. A practical first move is organisational. The IT teams that run storage and infrastructure cannot solve unstructured data quality and access on their own. They will need new partnerships across the business. Three partner groups matter most.
1. Partnering with security and compliance
Feeding internal data to AI pipelines raises immediate questions about what’s inside it. Sensitive records, regulated content, intellectual property, and personal data hide in file shares and object stores that no one has fully assessed or filtered. Security and compliance teams can classify sensitive data, determine which data types carry the most risk, and decide what to do with them once found, whether that means masking, quarantining, excluding, or moving them to controlled storage.
The barriers for IT are cultural as much as technical. Security and infrastructure teams often operate on different priorities and reporting lines and historically meet only when something breaks. Security may lack visibility into where unstructured data physically lives, while infrastructure may lack the context to judge what is sensitive.
A sensible path is to bring security into data decisions early rather than after an incident, agree on a shared classification scheme, and treat sensitive data discovery and mitigation as a continuous process, aided by automation, instead of a one-time audit. Joint ownership of a data inventory, with agreed policies for what happens when risky data is found, turns a reactive relationship into a proactive one.
2. Partnering with data engineers and analysts
The second partnership is with the data engineers and analysts who increasingly need unstructured data inside lakehouse platforms such as Snowflake and Databricks. These teams know what the models and pipelines require. They do not always know where the source data sits nor how to search it nor what it costs to move. Storage engineers hold that knowledge but may have limited understanding of how modern analytics tools consume data.
The barrier is a knowledge gap running in both directions. IT engineers and architects may not know how a lakehouse ingests, formats, or queries data, or why certain pipeline patterns are expensive. Data engineers may not appreciate the physics of moving petabytes between systems, but they want and may move unstructured data regardless of its state of preparation and quality anyway. Left unaddressed, this gap produces duplicated data, poor analytics results, surprise egress bills, and pipelines that break at scale.
To begin, storage engineers should build a working understanding of how analytics and AI platforms operate and what data they require, and data teams should understand the constraints and costs of the underlying storage. Shared planning on data curation and movement prevents expensive rework. Identifying preparation needs and problems before an AI project starts is far cheaper than discovering them in the middle of one.
3. Partnering with line of business heads
The third partnership is with the line of business leaders who understand what the data means. IT can see that a file share holds 10 million documents. Departments chime in with knowledge on which of those datasets are highest priority, what keywords people need to search for, what should be excluded, and what belongs in BI tools and AI models.
Even so, given the size of unstructured data and its overall lack of context, automation, such as with unstructured data management and AI data classification tools, is required to scan and enrich metadata so that it can be discovered via tags and keywords.
The soft skills requirement is translation. Business leaders think in outcomes and domain language, while IT thinks in systems and storage. A shared vocabulary for data classification and curation between these groups is valuable to get on the same page. This may entail jointly defining priority datasets, the search terms and metadata that matter, exclusion rules, and the criteria for what becomes available to analytics and AI.
Benefits of IT collaborations for AI data management
These partnerships can produce three things that are vital to the success of AI initiatives and broad adoption.
- The first is granular visibility, insight, and classification across every silo. You cannot govern, prioritise, or prepare data you cannot see.
- The second is lower costs. Cross-team collaboration surfaces the hard problems, including sensitive data exposure, expensive movement, and redundant copies, while they are still cheap to address rather than mid-project when they are not.
- The third is a clear method for both preparing and moving data: which data to clean and classify, which to leave in place, which to move, and how to do it without repeated, costly transfers.
How large companies are adapting
There is early evidence that large, progressive organisations are restructuring to support this work rather than bolting it onto existing roles. McKinsey’s research on AI adoption finds that the organisations capturing real value are markedly more likely to redesign workflows and operating models than to simply add tools, and that organisational change, not technology procurement, is the main differentiator.
On talent, the pattern leans toward reskilling existing staff and hiring more selectively. McKinsey reports that leading organisations are hiring fewer technologists overall while raising demand for senior engineers, architects, and others who can set standards and orchestrate work across internal teams, vendors, and AI agents, and that they are reskilling their existing workforces in parallel. Cross-functional teams and clear executive sponsorship are common threads.
None of this removes the need for the partnerships themselves. Reorganisation and reskilling make collaboration easier, but the core shift is simpler and harder at once. The teams that own infrastructure, security, analytics, and business strategies have to solve the unstructured data problem together, because none of them can solve it alone.
