Enterprise AI investments are growing fast, and so are expectations. Organizations are scaling deployments across the business with ambitious goals for productivity, efficiency, and growth.
But as enterprises move from experimentation to operational deployment, many are discovering that AI is more expensive and difficult to scale than expected. Recent HyperFRAME research found that 84% of surveyed enterprise infrastructure and operations leaders said AI deployment had consumed more budget and operational resources than originally planned.
A significant portion of the hidden cost of enterprise AI begins before a model ever reaches production. It starts with the work required to find, understand, prepare, govern, and continuously deliver the enterprise data that AI depends on. For many organizations, AI has exposed a data problem that was already there. The mistake is treating AI as a standalone technology initiative. Enterprise AI will ultimately be embedded across applications, operations, and business processes, which means its economics cannot be separated from the way an organization manages and modernizes its data estate. These challenges show up as real costs: scarce specialist time, duplicated integration work, manual validation, governance remediation, and the rework required when an AI pilot cannot scale beyond its original environment.
AI Does Not Eliminate Decades of Data Complexity
Large enterprises rarely have one clean, centralized source of information. Their data is distributed across mainframes, cloud platforms, SaaS applications, databases, and distributed systems accumulated over decades of technology investment.
However, some of the most valuable information is also some of the hardest to access. Mainframes, for example, continue to hold critical transactional data across industries such as banking, insurance, healthcare, retail, and transportation. These systems often contain high-value, business-critical transactional data – from financial activity and claims to inventory and customer records – that can be essential to enterprise AI use cases. That information can be enormously valuable to AI. It can also be difficult for modern data teams to discover and use.
The challenge is not simply moving data from one place to another. Teams need to understand what data exists, what it means, how different datasets relate to one another, who is permitted to access them, and whether the information being provided to an AI system is accurate and current.
Data Preparation Becomes an AI Infrastructure Problem
For years, employees filled in the gaps created by fragmented enterprise data. Over time, analysts, mainframe specialists, and application teams developed a working knowledge of where the right data lived, what different fields meant, and which datasets could be trusted. Much of that context is often rooted in years of experience and is not readily available to an AI model.
Making that context available to AI requires organizations to share that knowledge systematically throughout the organization. Making data available is only part of the task. AI also needs context: the business meaning, relationships, lineage, and policies that allow a model or agent to use that data correctly. Data discovery, metadata, lineage, integration, and governance all play a role, particularly when information spans mainframe, cloud, and distributed environments.
Predicting customer churn is a good example of how quickly this gets complicated. Payment history may live on the mainframe, while customer interactions sit in a cloud CRM and product usage and support records reside in other systems. Identifying which customers are likely to cancel requires bringing that information into a usable context for the model.
When only some of that information reaches the model, the resulting analysis is based on an incomplete view of the customer. Connecting those sources and making sure the data is understood and handled appropriately adds another layer of work to an AI initiative, and at enterprise scale, that work can quickly become more complex than the AI project itself.
Stale Data Creates Another Form of Cost
Access alone is not enough. AI also depends heavily on timing. A model making operational recommendations based on yesterday’s inventory, last week’s account status, or an incomplete view of a transaction can make a technically reasonable recommendation that is already wrong.
This is why AI readiness increasingly depends on how well organizations synchronize information across platforms.
Batch-oriented approaches may still make sense for some use cases. Others require data to reflect business activity as it happens. If a transaction changes on a core system, downstream analytics and AI applications may need to see that change quickly. Otherwise, organizations create a new kind of data silo: an AI environment that technically has access to enterprise information but does not have an accurate view of the enterprise right now.
When data is out of date, teams often end up doing more work around the AI system. They may check a recommendation against the source system before using it or spend time figuring out why two systems are showing different information. That extra work adds up, especially when AI is being used across multiple parts of the business.
Governance Becomes Harder as Data Starts Moving
AI also changes the data governance equation. Traditionally, governance focused heavily on where data was stored and who could access it. In an AI-driven organization, data increasingly moves across environments and into models, retrieval systems, automated workflows, and agents. Governance must account for which systems can use data, for what purposes and under what conditions, while ensuring sensitive information is protected and AI outputs remain traceable to the data that influenced them.
When those controls are applied inconsistently, organizations create additional work to validate access, trace data usage, and demonstrate that AI initiatives remain compliant as they scale.
Modernization Should Start with Access, Not Migration
One of the mistakes enterprises can make is assuming that AI readiness requires moving everything onto the newest platform when it does not. Organizations should absolutely modernize their infrastructure, but modernization does not have to mean relocating every system of record before its data becomes useful.
In many cases, the more practical strategy is to modernize access to the data first.
That means enabling modern analytics, cloud platforms, and AI systems to securely discover and work with information wherever it currently resides. It also means creating common context across systems, making metadata understandable, maintaining governance as information moves, and delivering current data without disrupting the applications that already run the business.
This approach changes the economics of AI modernization.
Instead of treating decades of infrastructure as technical debt that must be replaced before innovation can begin, organizations can preserve what already works while making its data available to new use cases.
The Foundation Determines the Return
As AI models become more widely available, the ability to deploy them will become less of a differentiator. The harder challenge will be giving those models a complete, current, and governed view of the business without creating more complexity in the process.
That puts data modernization directly in the path of AI ROI. Organizations that can efficiently make enterprise data understandable, current, and governed will be better positioned to scale AI without allowing operational costs to rise at the same rate.
Ultimately, AI investments depend on the data foundation beneath them. An accessible, current, and reliable foundation is what turns AI spending into business value.
Build your AI governance skills in 2026.
DATAVERSITY’s training programs cover AI governance, data governance, and compliance for data practitioners.
