Market map / Battle
Market map · scheduled battle

AI-native search and the post-Google information access market

September 13 at 20:47 · $0.094 total

claude-fable-5 wins
judged by gpt-5.6-luna-pro · both position-swapped passes agree
“A offers sharper value-chain segmentation, stronger economics and credible provenance white space; B misclassifies several products and makes vaguer white-space calls.”
claude-fable-5 · winner
$0.08227s1601 tokens

Market Map: AI-Native Search & Post-Google Information Access

1. Thesis

The unbundling of the "ten blue links" is shifting value from ad-monetized traffic aggregation to answer delivery embedded in workflows — collapsing the distinction between search, assistant, and agent, and forcing a rebuild of the entire economic layer (crawling rights, attribution, monetization) underneath.

---

2. Segments

A. Consumer Answer Engines

Perplexity, OpenAI (ChatGPT Search), Google (AI Overviews / Gemini), Microsoft (Copilot/Bing), You.com (You.com has pivoted increasingly toward enterprise/API — less sure of current consumer emphasis)

Dynamics: Brutal capital-intensity war; differentiation is thin because all rest on similar model capabilities and web indexes. Incumbents (Google) are cannibalizing themselves defensively. Monetization unsolved — CPM-style ads don't map cleanly onto single-answer surfaces.

B. Search Infrastructure & Retrieval APIs

Exa (formerly Metaphor), Brave (Search API), Tavily, Vespa.ai, Algolia, Elastic

Dynamics: "Selling shovels" — every agent and AI app needs web retrieval or semantic search primitives. Sticky, usage-based revenue, but exposed to hyperscalers bundling retrieval into model APIs (OpenAI/Anthropic native web search tools compress this layer).

C. Enterprise / Internal Knowledge Search

Glean, Hebbia, Sana (Sana — reasonably confident but verify current positioning), Coveo, Microsoft (Copilot for M365)

Dynamics: Highest willingness-to-pay in the market; ROI is legible (analyst/employee time). Permissions-aware retrieval across SaaS silos is genuinely hard, creating a real moat vs. generic LLMs. Microsoft's distribution is the existential threat.

D. Vertical / Domain-Specific Answer Engines

Harvey (legal), OpenEvidence (medical), Consensus (scientific literature), AlphaSense (financial/market intelligence), Elicit (research)

Dynamics: Depth beats breadth — proprietary corpora, citations that professionals can defend, and workflow integration justify premium pricing. Most defensible against foundation-model encroachment because trust and liability matter more than raw capability.

E. Content Economics & Attribution Layer

TollBit, Cloudflare (pay-per-crawl / bot management), ProRata.ai, Scale-adjacent licensing deals (e.g., OpenAI–News Corp, Perplexity Publisher Program)

Dynamics: Nascent and chaotic — the market is renegotiating who pays for the content answers are built on, via lawsuits (NYT v. OpenAI), licensing deals, and crawl-blocking. Whoever standardizes the transaction layer captures a toll position.

---

3. Where Value Accrues: Segment D (Vertical Answer Engines), with C close behind

  • Pricing power: Harvey and OpenEvidence charge for outcomes, not queries; no dependence on solving the ad-model problem that plagues Segment A.
  • Defensibility: Proprietary/licensed corpora + regulatory trust + workflow lock-in are moats general models can't easily replicate.
  • Segment A will see the most revenue flow through it but the worst value capture — margins are being competed away and Google/OpenAI will absorb the general case.
  • Segment B is real but compressible; Segment E is optionality with fat tails but binary legal/standards risk.

---

4. White Space

  1. Agent-to-agent commerce/search protocol layer: As agents (not humans) become the dominant "searchers," no player offers a neutral standard for machine-readable offers, structured intent, and agent-negotiated transactions — the "SEO + payments rail" for agentic traffic. (Stripe and Cloudflare circle this; no one owns it.)
  2. Provenance-verified answer infrastructure for regulated industries: A cryptographically auditable "chain of citation" — proving exactly which source versions produced an answer, at query time — for compliance, litigation, and audit contexts. Vertical players cite sources, but nobody sells verifiable citation infrastructure as a horizontal product.

---

Confidence notes: Company names and core positioning are current as of my knowledge cutoff (early 2025); funding stages and product pivots (especially You.com, Sana, ProRata) move fast and should be re-verified.

grok-4.6
$0.00722s1050 tokens

(1) Thesis Google’s “10 blue links + ads” monopoly is being unbundled into answer engines, agentic retrieval APIs, and query-owning surfaces (chat, browser, OS), so value migrates from CPC on links to subscriptions, usage-based infra, and workflow lock-in.

(2) Segments (real companies only)

Consumer answer engines / chat-search Perplexity, OpenAI (ChatGPT Search / Atlas-class products), Google (AI Overviews + Gemini), xAI (Grok), You.com, Brave Search. Dynamics: winner-take-most on default habit and freshness; still subsidized by ads or model-lab balance sheets; citation quality and “I can click through” trust are the choke points.

Query-owning browsers / OS layers The Browser Company (Arc / Dia), Brave, Opera (Aria), Apple (Safari + Apple Intelligence / Spotlight path), Microsoft (Edge + Copilot). Dynamics: whoever owns the address bar and default new-tab owns the query before any engine; distribution > model quality in the near term.

Web-scale retrieval / search APIs (for agents) Exa, Parallel, Tavily, Firecrawl, SerpAPI / Bright Data-class SERP, (less sure on relative scale) You.com API. Dynamics: agents need structured, citation-grade, low-latency web access that Google does not productize cleanly; this is becoming the “search index as a service” layer.

Enterprise / workplace search & RAG Glean, Microsoft Copilot (M365), Google Gemini for Workspace / Vertex Search, Notion AI, Elasticsearch / Elastic, Coveo. Dynamics: value is permissions, connectors, and “answer over our corpus,” not the open web; switching costs are high once connectors and ACLs are live.

Vertical / professional search Harvey / Casetext (legal), Consensus / Elicit / Semantic Scholar (science), GitHub Copilot / Sourcegraph (code), Bloomberg / AlphaSense (finance). Dynamics: willingness-to-pay is 10–100× consumer because of liability, recency, and domain corpus; incumbents with licensed data still have a moat.

(3) Who captures the most value Enterprise / workplace search (and, secondarily, vertical professional search). Reason: budgets are large and recurring, data gravity + ACL/connectors create lock-in, and the buyer is not an ad-supported consumer. Consumer answer engines will have huge usage but structurally weaker take rates (ads vs. Google, or thin subscriptions) unless they also own the browser/OS default.

(4) White space (no current player really covers)

  1. Private, lifelong personal memory search — a user-owned index of email, messages, photos, location, and browsing that answers “what did I already know / decide” without sending the corpus to a lab or ad network. Pieces exist (Apple on-device, Rewind-class, Notion) but not a full, portable, queryable personal graph.
  2. Physical-world / local multimodal search that is not ads-first — “what’s actually open, in stock, and true right now” from cameras, maps, and live sensors, with citations, not Yelp/Google Local ad ranking. Maps + LLMs have demos; no independent, trusted layer.