Market map / Battle
Market map · scheduled battle

The modern data stack after LLMs: where ETL, warehousing, and BI are heading

September 13 at 20:56 · $0.132 total

Judged a tie
judged by claude-fable-5 · passes split → scored as a tie
“B's value-flow argument (compute meter, bundling economics) is more rigorous and its placements accurate; A's semantic-layer value call and MetricFlow claim are shakier.”
deepseek-v4-pro
$0.01466s5256 tokens

Thesis: The modern data stack is being restructured from a batch ETL→warehouse→BI pipeline into an LLM-native, semantics-first control plane where data integration, storage, and analytics converge on governed, AI-ready context—metrics, lineage, vectors, and access policies—so that both human and agentic queries are trusted and traceable.

---

1. Ingestion, transformation & orchestration (ETL/ELT)

Companies: Fivetran, Airbyte, dbt Labs, Dagster, Prefect, Meltano Dynamics: LLMs are accelerating pipeline generation and schema mapping, but reliability, change data capture, and transformation testing remain the real moat; this layer is consolidating around a few open-source and managed standards.

2. Warehouse, lakehouse & vector engines

Companies: Snowflake, Databricks, Google BigQuery, ClickHouse, Pinecone, Weaviate Dynamics: Incumbents are absorbing vector search, RAG, and AI copilots directly into the lakehouse, forcing standalone vector DBs to compete on latency, scale, or hybrid search; data gravity still favors Snowflake and Databricks.

3. BI & analytics interfaces

Companies: Microsoft Power BI, Salesforce Tableau, Google Looker, ThoughtSpot, Hex, Metabase Dynamics: Dashboards are being augmented or replaced by conversational and agentic analytics; incumbents have distribution, but AI-native tools are winning data-science and product teams with faster natural-language and notebook-based workflows.

4. Semantic, metrics & LLM context layer

Companies: dbt Labs (MetricFlow), Cube, Atlan, Alation, Glean, LlamaIndex (less sure: LlamaIndex enterprise revenue) Dynamics: This is becoming the new control point: semantic definitions, lineage, retrieval context, and access policies for AI answers; the space is fragmented across catalogs, metric layers, and vector/retrieval tools, with consolidation likely.

5. Governance, quality & observability

Companies: Monte Carlo, Soda, Bigeye, Datafold, Collibra, Great Expectations Dynamics: LLM-era data issues shift from “table quality” to “answer trust”; observability must extend to RAG retrieval, embedding drift, and LLM outputs, but most existing tools still monitor structured tables and dashboards.

---

Which segment wins the most value and why

The semantic, metrics & LLM context layer wins the most incremental value. Whoever owns the mapping from business data to metrics, embeddings, lineage, and access policies becomes the trust layer for every AI answer. Without that layer, LLMs hallucinate or leak data; with it, every query—human or agentic—is governed. Snowflake and Databricks will try to absorb this layer, but a warehouse-agnostic semantic system of record is the highest-leverage position.

---

Open white-space opportunities

  1. End-to-end LLM data contracts and drift monitoring
  2. No current player covers the full chain from source schema to embedding/vector drift to RAG output with automated retraining, alerting, and audit. Existing observability tools monitor structured tables; vector databases do not provide governance or drift detection for LLM context.
  3. Universal agentic semantic control plane
  4. No one provides a warehouse-agnostic semantic layer that enforces row/column-level security across Snowflake, Databricks, Postgres, and SaaS apps, while serving both human BI and AI agents. dbt MetricFlow is warehouse-specific, Atlan is catalog-centric, and Cube is developer-oriented—none owns the end-to-end agentic governance plane.
gpt-5.6-luna-pro
$0.01027s6813 tokens

1. Thesis

LLMs are restructuring the modern data stack from a collection of human-operated pipelines and dashboards into an AI-accessible data control plane, shifting value toward governed semantics, unified storage/compute, and systems that can safely answer questions or take actions across the stack.

2. Market map

A. Data ingestion, integration, and replication

Companies: Fivetran, Airbyte, Informatica, Rivery, Estuary

Dynamics: Basic connector-based ETL is becoming commoditized, while differentiation is moving toward real-time replication, unstructured and application data, reverse ETL, data contracts, and reliable movement into multiple warehouses and AI systems.

  • Fivetran — managed ELT and CDC, with broad enterprise adoption.
  • Airbyte — open-source and commercial connector platform; competing on extensibility and cost.
  • Informatica — incumbent enterprise integration, MDM, and data governance platform.
  • Rivery — cloud-native data integration and orchestration.
  • Estuary — real-time data movement and CDC.

B. Transformation, orchestration, and data quality

Companies: dbt Labs, Dagster, Prefect, Astronomer, Coalesce

Dynamics: SQL-based transformation remains central, but pipelines are evolving into software products with version control, tests, lineage, observability, and AI-assisted development; orchestration vendors are competing to become the operating layer for both batch data and AI workflows.

  • dbt Labs — the leading analytics engineering and SQL transformation ecosystem.
  • Dagster — developer-oriented orchestration with strong asset-based abstractions.
  • Prefect — Python-first workflow orchestration.
  • Astronomer — managed Apache Airflow and enterprise orchestration.
  • Coalesce — visual and metadata-driven data transformation, particularly around Snowflake.

C. Cloud data warehouses, lakehouses, and unified data platforms

Companies: Snowflake, Databricks, Google BigQuery, Microsoft Fabric, Amazon Redshift

Dynamics: Warehouses and lakehouses are converging into broader data platforms that combine storage, compute, governance, streaming, notebooks, ML, and increasingly native AI capabilities; platform consolidation is a major threat to point tools.

  • Snowflake — cloud data warehouse expanding into applications, governance, AI, and transactional workloads.
  • Databricks — lakehouse platform with strong positions in engineering, ML, and increasingly SQL analytics.
  • Google BigQuery — serverless warehouse tightly integrated with Google Cloud and Gemini.
  • Microsoft Fabric — integrated warehouse, lake, BI, engineering, and AI platform built around the Microsoft ecosystem.
  • Amazon Redshift — AWS’s warehouse platform, increasingly integrated with the broader AWS analytics stack.

D. Metadata, governance, security, and semantic control

Companies: Collibra, Alation, Atlan, Cube, GoodData

Dynamics: This layer is becoming more important as LLMs need trustworthy definitions, permissions, lineage, and business context; however, the market remains fragmented between catalog/governance products and narrower semantic or metrics layers.

  • Collibra — enterprise data governance, cataloging, and stewardship.
  • Alation — data catalog and knowledge-sharing platform.
  • Atlan — modern metadata, catalog, lineage, and collaboration layer.
  • Cube — headless BI and semantic modeling infrastructure.
  • GoodData — embedded analytics and semantic-layer technology.

Less certain boundary: Semantic-layer vendors are increasingly being absorbed into larger platforms or embedded inside BI and transformation tools, so this may become a capability rather than a standalone category.

E. BI, augmented analytics, and natural-language data applications

Companies: Microsoft Power BI, Tableau, Looker, ThoughtSpot, Sigma Computing, Hex

Dynamics: Traditional dashboarding is mature and increasingly bundled, while growth is shifting toward governed natural-language exploration, spreadsheet-like interfaces, embedded analytics, and AI agents that can explain metrics or perform analysis.

  • Power BI — dominant enterprise BI distribution through Microsoft.
  • Tableau — large installed base and strong visualization brand under Salesforce.
  • Looker — model-centric BI tightly integrated with Google Cloud.
  • ThoughtSpot — search- and AI-led analytics.
  • Sigma Computing — cloud-native, spreadsheet-style analytics.
  • Hex — collaborative notebooks combining SQL, Python, and presentation.

3. Segment likely to capture the most value

Cloud data platforms and unified data control planes

This segment is likely to capture the most aggregate economic value because:

  1. They control the primary consumption meter. Storage, compute, query execution, and model-serving workloads generate recurring infrastructure spend.
  2. LLMs increase—not reduce—the need for data compute. AI applications require more data preparation, retrieval, feature computation, evaluation, and inference-adjacent workloads.
  3. They can bundle adjacent functions. Snowflake, Databricks, BigQuery, and Fabric can progressively absorb ingestion, governance, transformation, vector search, ML, and BI capabilities.
  4. They own strategic distribution. CIOs generally prefer fewer platforms with integrated security, procurement, and support.
  5. They can monetize both the data plane and AI plane. The winning platforms will charge for storage, compute, data sharing, model access, and application workloads.

The semantic and governance layer may capture disproportionate strategic influence, particularly if it becomes the trusted interface for AI agents. But its standalone revenue pool is likely to be smaller unless it controls execution, policy, or a high-volume operational workflow.

4. White-space opportunities

1. A cross-stack, governed analytics agent

Opportunity: An agent that can inspect metadata, understand business definitions, generate or modify transformation code, run queries, trace results to source data, explain uncertainty, and request approval before making production changes—across Snowflake, Databricks, BigQuery, dbt, BI tools, and governance systems.

Why it is open: Current products typically optimize one layer: BI copilots answer questions, dbt assists with transformations, catalogs provide context, and warehouses provide execution. No broadly adopted product appears to own the complete closed loop from question → semantic interpretation → data preparation → governed query → verified answer → reusable asset.

Potential buyer: Data platform teams and business operations in large enterprises.

2. An autonomous data reliability and cost-control layer

Opportunity: A cross-platform system that continuously detects broken pipelines, schema changes, stale or duplicated data, metric drift, poor query economics, unnecessary warehouse spend, and low-quality AI retrieval contexts—and then proposes or safely executes fixes.

Why it is open: Observability vendors monitor subsets of the stack, while cloud platforms optimize their own environments. There is still no dominant neutral layer that jointly optimizes data correctness, freshness, lineage, AI answer quality, security policy, and infrastructure cost across multiple clouds and tools.

Potential buyer: Enterprises operating heterogeneous stacks after years of tool adoption and cloud migration.

Less certain: Monte Carlo, Bigeye, Datadog, cloud-native observability products, and the major data platforms all address portions of this opportunity. The white space is the cross-vendor, closed-loop version rather than basic monitoring.