The modern data stack after LLMs: where ETL, warehousing, and BI are heading
September 13 at 20:42 · $0.129 total
The Modern Data Stack After LLMs: Market Map
1. Thesis
LLMs are collapsing the modern data stack's unbundled layers back toward consolidated platforms, shifting value from tools that move and store data to systems that make data AI-ready and queryable in natural language — turning the warehouse into the context layer for enterprise AI rather than the endpoint for dashboards.
---
2. Segments
A. Warehouse/Lakehouse Platforms (the consolidators)
Snowflake, Databricks, Google BigQuery, Microsoft Fabric, AWS (Redshift/S3 + Bedrock)
Dynamics: The gravity centers. Both Snowflake (Cortex) and Databricks (Mosaic acquisition, DBRX) are racing to become the "AI + data" platform, absorbing adjacent layers (ingestion, transformation, serving). Databricks' acquisitions (MosaicML, Tabular, Neon) signal a full-stack land grab. Open table formats (Iceberg, Delta) are commoditizing storage, pushing competition up to compute + AI services.
B. Ingestion / ETL / Transformation (the squeezed middle)
Fivetran, Airbyte, dbt Labs, Matillion, Informatica (acquired by Salesforce, 2025)
Dynamics: Under existential pressure. Connectors are being commoditized (LLMs can generate/maintain them; warehouses bundle ingestion — e.g., Snowflake's Openflow). dbt is pivoting toward being the semantic/metadata layer for AI rather than just transformation. Expect consolidation — Fivetran's merger with dbt Labs (announced 2025; I'm fairly confident but verify status) exemplifies defensive bundling.
C. BI → Conversational / Agentic Analytics
Tableau (Salesforce), Power BI (Microsoft), Looker (Google), ThoughtSpot, Hex, Sigma Computing
Dynamics: The most visibly disrupted layer. "Chat with your data" threatens the dashboard paradigm; incumbents are bolting on copilots (Tableau Pulse, Power BI Copilot) while natives like ThoughtSpot and Hex rebuild around NL-first workflows. The hard problem isn't the LLM — it's the semantic layer that makes NL queries correct, which is why BI value is migrating downward into metrics/semantic layers (dbt Semantic Layer, Cube).
D. AI-Native Data Infrastructure (the new layer)
Pinecone, Weaviate, LangChain (LangSmith), Unstructured.io, LlamaIndex, Glean
Dynamics: Vector DBs, unstructured-data ETL, and retrieval orchestration — the "modern data stack for LLMs." Standalone vector DBs are getting squeezed as Postgres (pgvector), Elastic, MongoDB, and the warehouses add vector search natively; value is shifting to unstructured-data processing (Unstructured.io) and enterprise retrieval/context (Glean).
E. Metadata, Governance & Semantic Layer
Atlan, Collibra, Alation, Monte Carlo, Cube
Dynamics: Previously a compliance afterthought, now strategically central: AI agents need trustworthy lineage, definitions, and quality signals to answer correctly. Data observability (Monte Carlo) and catalogs (Atlan) are repositioning as the "context/trust layer for AI." Risk: warehouses build this in (Snowflake Horizon, Databricks Unity Catalog).
---
3. Where the Value Accrues
Segment A (warehouse/lakehouse platforms) wins the most value, for three reasons:
- Data gravity + workload capture: They already hold governed enterprise data; running AI where data lives beats moving data to AI. Every layer they absorb (ETL, vector search, semantic models, even BI via Fabric/Cortex Analyst) is incremental revenue on an existing footprint.
- The squeeze works in both directions: ETL below and BI above are being commoditized/absorbed; the platform in the middle captures the margin.
- Consumption pricing on AI workloads: Agentic queries multiply compute consumption — the platforms are billed per use, so more AI activity = direct revenue.
The dark horse for relative value creation is Segment E — the semantic/context layer is the actual bottleneck to accurate enterprise AI — but it's structurally at risk of platform absorption, so much of that value may accrue to Snowflake/Databricks anyway.
---
4. White-Space Opportunities
1. Agent-native data contracts & permissioning ("IAM for AI agents over data") No one owns fine-grained, auditable authorization for autonomous agents querying enterprise data — row/column-level policy that travels with context into prompts, with usage attribution and cost governance per agent. Governance players (Immuta comes closest) are human-user-centric; observability tools watch pipelines, not agents. Whoever solves "which agent can see/derive what, and prove it" becomes mandatory infrastructure for regulated-industry AI adoption.
2. Verified semantic ground truth as a product — a "correctness layer" for NL analytics Semantic layers (dbt, Cube) define metrics, but nothing exists that continuously tests whether an LLM's answer over enterprise data is right: benchmark generation from a company's own schema, regression testing of text-to-SQL/agents against certified answers, and confidence scoring exposed to end users. It's the analytics equivalent of eval infrastructure — currently done ad hoc in every enterprise, owned by no vendor. (Adjacent players like Braintrust do general LLM evals, but none are data-stack-native — lower confidence on how fast this niche gets crowded.)
1) Thesis
LLMs are restructuring the modern data stack from a modular pipeline—ingest, warehouse, transform, dashboard—into an integrated, governed “data-to-decision” platform in which the warehouse/lakehouse becomes the execution layer and semantic context, agents, and operational actions become the new points of differentiation.
2) Market map
| Segment | Representative companies | Segment dynamics | |---|---|---| | 1. Data ingestion, replication, and activation | Fivetran, Airbyte, Qlik, Matillion, Census, Hightouch | Ingestion is becoming increasingly commoditized, while value is moving toward real-time replication, reverse ETL, APIs, and activation into CRM, marketing, and operational systems. AI-generated connectors may further pressure pricing, although reliability and governance remain defensible. | | 2. Warehousing, lakehouses, and data infrastructure | Snowflake, Databricks, Google BigQuery, Amazon Redshift, Microsoft Fabric, ClickHouse | This is the gravitational center of the stack. The distinction between warehouse, lakehouse, streaming platform, vector database, and AI infrastructure is blurring as vendors compete to own storage, compute, governance, and model-serving workloads. | | 3. Transformation, orchestration, and data operations | dbt Labs, Dagster, Prefect, Astronomer, Matillion, Coalesce | SQL transformation and workflow orchestration are converging into software-development-style data engineering. LLMs can generate pipelines, but production deployment still requires testing, lineage, dependency management, backfills, and incident response. | | 4. Semantic layer, BI, and AI-native analytics | Microsoft Power BI, Salesforce Tableau, Google Looker, Sigma Computing, ThoughtSpot, Hex | Traditional dashboards are being supplemented by natural-language querying, embedded analytics, notebooks, and agentic analysis. The strategic question is whether the semantic layer is owned by the BI vendor, the warehouse, or an independent metrics platform. | | 5. Governance, catalog, quality, and security | Collibra, Alation, Atlan, Monte Carlo, BigID, Soda | AI increases the value of lineage, access control, data contracts, quality monitoring, and sensitive-data discovery. Governance is shifting from passive cataloging toward policy enforcement at query, model, and agent-action time; procurement cycles remain long and fragmented. |
How the stack is changing
- ETL becomes ELT plus activation: Data is increasingly landed in the central platform, transformed there, and then pushed back into business applications.
- Warehouses become AI runtimes: They are absorbing vector search, unstructured data, model orchestration, streaming, and application development.
- BI becomes conversational and operational: The output is no longer only a chart; it may be a recommendation, workflow, or automated action.
- The semantic layer becomes strategic: Reliable definitions of metrics, entities, permissions, and business context are essential for preventing LLMs from producing plausible but incorrect answers.
- Governance moves upstream and into execution: Quality and access policies must apply to pipelines, models, prompts, generated SQL, and agent actions—not just reports.
3) Which segment captures the most value?
Winner: Warehousing, lakehouses, and data infrastructure
This segment is most likely to capture the largest absolute share of value because it controls the platform on which nearly every other workload runs.
Why:
- Consumption-based economics: Storage, compute, streaming, and AI workloads generate recurring usage revenue.
- High switching costs: Once data, governance, pipelines, and applications are deployed on a platform, migration is expensive.
- Control of AI workloads: LLM applications need governed enterprise data, vector search, inference orchestration, and low-latency access—capabilities increasingly bundled into the data platform.
- Platform consolidation: Snowflake, Databricks, Microsoft, Google, and AWS can bundle adjacent ingestion, transformation, governance, and BI functionality.
- Budget ownership: These platforms often sit within large data, infrastructure, or cloud budgets rather than narrower analytics-tool budgets.
The likely outcome is high absolute value capture by the platform layer, but attractive venture opportunities in control points around it—especially governance, semantic context, reliability, and workflow automation. BI may capture substantial user attention, but it is more exposed to bundling by Microsoft, Google, Salesforce, and the warehouse vendors.
4) Two white-space opportunities
Opportunity 1: A truly agent-ready enterprise semantic and policy layer
Problem: Existing catalogs, metrics layers, and BI semantic models usually describe tables and metrics, but do not fully provide the context an agent needs: business definitions, entity relationships, approved actions, permissions, source confidence, exception handling, and auditability.
Potential product: A vendor-neutral semantic control plane that:
- Unifies structured data, documents, events, and business processes
- Maintains canonical entities and metric definitions
- Generates governed context for LLMs and agents
- Validates generated SQL and proposed actions
- Enforces row-, column-, and action-level permissions
- Records provenance and explanations for every answer or action
Why it is white space: No current player appears to cover this end to end. Atlan, Collibra, dbt, Looker, Snowflake, Databricks, and specialist metrics vendors each cover important pieces, but the full combination of semantic modeling, agent context, policy enforcement, and action governance remains fragmented. Confidence: high that the category is underserved; lower that it is entirely unoccupied.
Opportunity 2: An autonomous, cross-stack data reliability and cost-control plane
Problem: Enterprises operate pipelines across multiple clouds, warehouses, SaaS sources, orchestration tools, transformation frameworks, and AI applications. Existing observability products often monitor data quality or pipeline health, but not the full causal chain from source change to broken dashboard, model, agent response, or cloud bill.
Potential product: A control plane that:
- Builds a live dependency graph across ingestion, transformation, warehouse, BI, and AI systems
- Detects schema, freshness, quality, semantic, and model-serving failures
- Identifies the business impact of an incident
- Automatically proposes or executes repairs
- Manages warehouse compute and pipeline costs
- Enforces data contracts and service-level objectives across vendors
Why it is white space: No incumbent clearly provides comprehensive cross-stack reliability plus automated remediation and FinOps. Monte Carlo, Bigeye, Soda, Datadog, dbt, and cloud-native tools cover subsets, but the market remains fragmented. Confidence: medium; this is an underserved integration layer rather than a completely absent product category.