Market map / Battle
Market map · scheduled battle

The modern data stack after LLMs: where ETL, warehousing, and BI are heading

September 13 at 21:11 · $0.097 total

grok-4.6 wins
judged by claude-fable-5 · both position-swapped passes agree
“B has cleaner segmentation (5 well-carved layers incl. governance), more accurately placed companies with staging caveats, and equally sharp value/white-space calls.”
gemini-3.1-pro
$0.01815s1519 tokens

Here is a venture analyst’s end-to-end market map of the Modern Data Stack (MDS) in the era of LLMs.

1. Thesis

LLMs are restructuring the modern data stack from a deterministic, pipeline-to-dashboard assembly line into an agentic, conversational ecosystem where unstructured data is a first-class citizen and insights are generated on demand rather than pre-computed.

2. Market Segments & Dynamics

Segment 1: AI-Native ETL & Unstructured Data Ingestion

  • Companies: Unstructured, Airbyte, Fivetran, dbt Labs, Snorkel AI.
  • Dynamics: Traditional ETL players are racing to add vectorization and unstructured data parsing, while AI-native upstarts (like Unstructured) are building pipelines specifically designed to feed complex documents into LLMs and vector databases.

Segment 2: Vector-Augmented Warehouses & Lakehouses (Storage & Compute)

  • Companies: Databricks, Snowflake, Pinecone, Weaviate, ClickHouse.
  • Dynamics: The battle here is between dedicated vector databases (Pinecone, Weaviate) and incumbent data warehouses (Databricks, Snowflake) that are aggressively absorbing vector search capabilities to prevent data from leaving their ecosystems.

Segment 3: The Generative BI & Semantic Layer

  • Companies: ThoughtSpot, Defog.ai, Vanna.ai, Cube, dbt Labs. (Note: Defog and Vanna are earlier-stage startups, but represent the pure-play Text-to-SQL movement).
  • Dynamics: Traditional static dashboards are being replaced by conversational interfaces and "chat-to-data" tools, forcing a massive reliance on robust Semantic Layers (Cube, dbt) to prevent LLMs from hallucinating SQL queries.

Segment 4: Data Quality & LLMOps for the Enterprise

  • Companies: Arize, Braintrust, Patronus AI, Monte Carlo, Anomalo.
  • Dynamics: Traditional data observability (Monte Carlo) is colliding with LLM evaluation (Braintrust, Patronus); enterprises cannot deploy generative BI without automated guardrails to verify that the LLM's data outputs are accurate and compliant.

3. Which Segment Wins the Most Value and Why?

Vector-Augmented Warehouses & Lakehouses (Databricks, Snowflake) will capture the most value. Why: Data gravity. Moving petabytes of data is expensive and introduces security/compliance risks. Rather than piping data out to external LLM tools or standalone vector databases, enterprises prefer to bring the compute (LLMs) to the data. By offering native vector storage, fine-tuning, and hosted LLMs (e.g., Databricks acquiring MosaicML, Snowflake's Cortex), the foundational storage layers will successfully bundle and commoditize many of the standalone AI tools built around them.

4. White-Space Opportunities

  • Opportunity 1: Automated Semantic Layer Generation. Currently, text-to-SQL fails because LLMs don't understand the messy, undocumented business logic of a company's raw tables. There is a white space for an AI agent that crawls a company's dbt repos, historical SQL queries, and Slack channels to automatically build and maintain a perfect, machine-readable semantic layer (ontology) without requiring months of manual data engineering.
  • Opportunity 2: Agentic Reverse-ETL. Traditional Reverse-ETL (e.g., Hightouch, Census) simply syncs warehouse data back into SaaS apps (like Salesforce). The white space is an "Agentic" Reverse-ETL platform where LLMs don't just sync data, but autonomously execute workflows based on warehouse changes (e.g., "The warehouse shows a user's usage dropped 20%; automatically draft a personalized email, adjust their Salesforce lead score, and create a Zendesk ticket for the CSM").
grok-4.6 · winner
$0.00929s1390 tokens

Thesis: LLMs are collapsing the classic MDS (rigid ELT → warehouse → dashboard) into AI-orchestrated, semantic, and agentic systems that treat data as queryable context rather than pipelines-to-be-maintained, so value shifts from specialized connectors/dashboards toward platforms that own compute + semantics + natural-language interfaces.

Segments

  1. AI-augmented ELT / reverse ETL & orchestration
  2. Fivetran, Airbyte, dbt Labs, Hightouch, Census, Dagster. Dynamics: Batch connectors persist but LLMs now auto-generate mappings, tests, and incremental models; reverse ETL is being absorbed into activation layers as warehouses add native outbound.
  3. AI-native warehouses / lakehouses
  4. Snowflake (Cortex), Databricks (Mosaic AI), Google BigQuery + Gemini, AWS (Redshift + Bedrock), MotherDuck. Dynamics: Storage/compute gravity plus vector indexes, LLM functions, and governance make these the default “context engines”; they are racing to become the operating system rather than just SQL engines.
  5. Semantic / metrics layers
  6. dbt Semantic Layer (MetricFlow), Cube, AtScale, GoodData. Dynamics: The missing contract between raw tables and LLMs; whoever owns consistent metrics and business definitions becomes the control plane for both humans and agents. (Transform was acquired; coverage here is thinner than the first two segments.)
  7. Conversational / agentic analytics & BI
  8. ThoughtSpot, Hex, Sigma Computing, Tableau Pulse / Einstein, Microsoft Power BI Copilot. Dynamics: Dashboards are being replaced by NL-to-SQL + viz + follow-up agents; incumbents bolt Copilots on while specialists compete on reliability and governed semantics. Emerging names (Seek AI, Defog, Vanna) are still small.
  9. Observability, catalog & governance with AI
  10. Monte Carlo, Atlan, Alation, Collibra, Bigeye. Dynamics: LLMs surface lineage, anomalies, and PII automatically, but the category remains a tax on the stack rather than a growth engine unless it becomes the policy layer for agents.

Value capture: Intelligent warehouses/lakehouses (segment 2) win the most. They already sit on the data, own the compute bill, and can natively add vectors, LLM functions, and semantic APIs; every other layer either feeds them or queries them. Switching costs and usage-based pricing compound this.

White-space (no current player fully covers)

  • End-to-end autonomous data-engineering agents that write, test, deploy, and self-heal pipelines from a ticket or Slack message with production-grade reliability and audit trails.
  • A unified structured + unstructured semantic graph (tables + docs + Slack + tickets) that LLMs can query with guaranteed metric consistency and row-level security—today’s catalogs and warehouses still treat these as separate worlds.