Enterprise leaders are under growing pressure to turn AI experimentation into measurable business outcomes. With the focus on how the technology can improve productivity, automate processes, or unlock new revenue streams, many have already launched pilots and proof-of-concepts.
Many of these initiatives, however, struggle to move into production. A recurring obstacle is the lack of data that is sufficiently reliable, accessible and governed to support AI at scale. Gartner expects organizations to abandon 60% of AI projects unsupported by AI-ready data in 2026.
Before an AI model can deliver reliable results, it needs access to trusted data.
Why Data Platforms Suddenly Matter Again
Customer data lives in the CRM, financials reside in a separate ledger, and operational data sits in the ERP. Everyone is generating new reports based on data on their laptops, and with the cloud, it is easy to become a citizen data engineer – creating a new silo has never been easier.
Adding AI to the mix only amplifies the scenario. Garbage In/Garbage Out still holds true, except that unbridled AI can turn that garbage into confident, wrong answers at scale.
For years, data platforms were treated largely as migration projects: systems to be moved, modernized, and then left alone until the next migration cycle. Now they are the core of any enterprise AI use case. Data engineering, analytics, governance and AI are all part of the same ecosystem. The priority is shifting from disconnected platforms for warehousing, analytics and machine learning towards a more unified data foundation, with governance built into the architecture.
Data lakes, meant to fix rigid warehouses, often became data swamps, while clients kept legacy integrations live alongside them. Modern lakehouse architectures have emerged as a response, giving organizations one common foundation for analytics, data engineering, governance and AI.
This is one reason why platforms such as Databricks have gained so much attention. Open table formats such as Delta Lake support transactional consistency and versioned data, helping organizations maintain reliable table histories and build more robust foundations for governed data operations.
Governance Moves to the Center
A common data foundation solves only part of the problem. The next question is whether the data can be trusted, traced and governed.
Until a few years ago, data lineage, access controls and catalogs were primarily concerns for governance teams, implemented as an afterthought. Today, governance impacts enterprise goals and AI outcomes. If an organization cannot explain where data came from, who modified it, how it was transformed and who has access to it, it cannot trust the AI outcomes built on top of it.
Unity Catalog provides a unified framework for managing access, lineage and auditing across the lakehouse. These capabilities can help organizations strengthen the data governance and traceability needed to meet applicable regulatory obligations, including those arising under the EU AI Act.
Scaling across Industrial Enterprises
Industrial businesses today increasingly need to combine traditional IT data with inputs from machines, sensors, production assets, and operational systems. The volumes here are enormous – a plant with 5,000 sensors sampling once a second produces over 400 million readings a day.
Till now this information was often siloed onsite at the plant – whether in Bangalore or Johannesburg – while the corporate office generated reports based on Excel files shared over email. Moving the data seamlessly off the plant floor, however, continues to pose a challenge.
Ingestion technologies such as Auto Loader and Zerobus can help bring plant-level data into a governed data environment, creating a foundation for applications such as predictive maintenance, production optimization, supply-chain visibility, asset performance management and energy optimization. And while these outcomes are often described as AI engagements, they are clear data initiatives first.
AI Is an Outcome, Not a Strategy
One of the biggest mistakes organizations make is treating AI as the starting point.
In my experience, however, the strongest results are derived from a focus first on creating trusted data foundations, involving a modernization of architecture, improvement of governance, and investments in data engineering and observability. While this may not sound as exciting as talking about autonomous agents and intelligent assistants, it is usually what separates successful transformations from endless POCs.
This does not mean fixing every dataset first. We need to build the governed foundation one high-value use case at a time and let each use case extend it.
What needs to be remembered in context is that AI creates value from data. And if the data cannot be trusted, governed and made accessible, no amount of AI sophistication can compensate for it.
So, before your next AI steering committee, ask one question: “If this model gave a wrong answer tomorrow, could we trace why?"
If not, you have a data problem.