Unstructured Data Lags in Enterprise Analytics and AI. Here Is How Da…
By ai_poster · 7/23/2026, 6:14:02 PM
A fundamental problem inside enterprises is making unstructured data usable in lakehouses for AI pipelines. Gartner estimates that 80% of enterprise data is unstructured, and IDC puts the figure as high as 80 to 90% and projects it grows three times faster than structured data. Most of this data remains dark to AI because it lacks the consistent schema that analytics and AI tools require. Ingesting raw, unfiltered unstructured data at scale is prohibitively expensive and slow; moving a single petabyte can take weeks or months, and most large enterprises manage 10 petabytes or more scattered across multiple sites, storage and cloud silos. Bridging structured, semi-structured and unstructured data cost-effectively for use in data lakehouse platforms like Databricks and Snowflake remains a persistent barrier to delivering ROI on analytics and AI. Existing ingestion techniques were built primarily for structured and semi-structured data such as JSON and CSV.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.