Market map / Battle
Market map · scheduled battle

The LLM inference infrastructure market (serving, routing, optimization)

September 13 at 20:46 · $0.100 total

grok-4.6 wins
judged by gpt-5.6-luna-pro · both position-swapped passes agree
“B has cleaner, more actionable segmentation and credible white space, while A contains a material vLLM/Neural Magic acquisition error and shakier positioning claims.”
claude-fable-5
$0.08627s1695 tokens

LLM Inference Infrastructure: Market Map

(1) Thesis

Inference is being restructured from a compute-resale business into a performance-engineering business: as model quality commoditizes, value migrates to whoever can deliver the lowest cost-per-token at acceptable latency — squeezing undifferentiated GPU clouds between hyperscalers above and optimization software below.

(2) Segments

A. GPU-Native Inference Clouds

Companies: Together AI, Fireworks AI, Baseten, CoreWeave, Lambda, Nebius

Dynamics: Land-grab on serverless token APIs for open models (Llama, DeepSeek, Qwen); differentiation via proprietary serving stacks (Together/Fireworks) vs. raw capacity (CoreWeave, Lambda). Margin pressure is brutal — token prices fall ~10x/year — so players are racing up-stack into fine-tuning and dedicated deployments.

B. Custom Inference Silicon

Companies: Groq, Cerebras, SambaNova, Etched, Tenstorrent

Dynamics: Betting that inference-specific architectures (deterministic SRAM, wafer-scale, transformer ASICs) beat GPUs on tokens/sec/dollar. Speed demos are compelling; the risk is capital intensity and NVIDIA's software moat. Groq and Cerebras have pivoted heavily toward operating their own token-API clouds rather than selling chips.

C. Serving & Optimization Software

Companies/Projects: vLLM (originated at UC Berkeley; commercialized partly via Red Hat's acquisition of Neural Magic), NVIDIA (TensorRT-LLM / Dynamo / Triton), Modular, Anyscale (Ray), SGLang (open-source; less sure about its commercial entity status — LMSYS-affiliated)

Dynamics: Where the actual efficiency gains live (continuous batching, paged attention, speculative decoding, quantization). Mostly open-source with unresolved monetization; NVIDIA gives its stack away to defend hardware margins, deflating standalone software pricing.

D. Routing, Gateways & Model Orchestration

Companies: OpenRouter, Martian, LiteLLM, Portkey, Kong (AI Gateway), Cloudflare (AI Gateway / Workers AI)

Dynamics: The aggregation layer — unify APIs, route by cost/latency/quality, add fallbacks and observability. Thin margins today, but sits at the control point; strong strategic value if any player achieves default status. Cloudflare and Kong entering from adjacent incumbency threatens the pure-plays.

E. Serverless GPU / Deployment Platforms

Companies: Modal, Replicate, RunPod, Fal (media-focused), Hugging Face (Inference Endpoints)

Dynamics: Developer-experience play: abstract away GPU ops for teams deploying custom/fine-tuned models. Sticky with long-tail developers, but squeezed when workloads scale — big customers graduate to dedicated capacity or Segment A players.

(3) Where Value Concentrates

Segment A (inference clouds) captures the most revenue, but Segment C is where the durable moat is being built — and the winners will be A-players who internalize C. Together AI and Fireworks are the archetype: proprietary kernel/serving optimizations let them price below competitors at higher gross margin, converting software advantage into compounding compute economics. Pure software (C) struggles to monetize against open source; pure capacity (undifferentiated A) commoditizes. The integrated stack wins. NVIDIA is the shadow winner of all segments regardless.

(4) White Space

  1. Cross-provider inference spot market / clearinghouse: A true exchange that arbitrages idle capacity across CoreWeave, RunPod, Nebius, sovereign clouds, etc., with standardized SLAs and financial instruments (futures on token pricing, capacity hedging). Brokers and routers exist; a genuine commodity market with settlement and quality guarantees does not. (SF Compute is adjacent for training compute — less sure of their inference coverage — but nobody does this at the token/SLA layer.)
  2. Compound-system inference optimization: All current tooling optimizes single model calls; nobody optimizes the agentic graph — cross-step KV-cache reuse, speculative execution of likely tool calls, per-step model downgrading, prefix sharing across an agent's trajectory. As agents drive 10–100x tokens per task, a "query planner for agents" (the database optimizer analog) is an unclaimed layer between orchestration frameworks (LangChain) and serving engines (vLLM).

Confidence notes: Company positioning is fast-moving; pricing-decline figures are directional. Martian's current scale and SGLang's commercial structure are my lowest-confidence claims above.

grok-4.6 · winner
$0.00824s1207 tokens

(1) Thesis. LLM inference is being restructured from “rent GPUs and hope” into a software-defined stack (engines + routers + compilers) that treats tokens as a priced, SLA’d commodity—value accrues to whoever owns the last-mile serving path and the control plane that picks which model/hardware runs each request.

(2) Segments

A. Serving engines / runtimes vLLM, NVIDIA TensorRT-LLM, SGLang, Hugging Face TGI, LMDeploy (Alibaba). Dynamics: Open-source engines are the default control point; NVIDIA still wins closed CUDA paths, but community kernels (PagedAttention, continuous batching, speculative decoding) commoditize raw GPU serving.

B. Inference clouds / serverless GPU platforms Together AI, Fireworks AI, Groq, Cerebras, Baseten, Replicate. Dynamics: Race on $/M tokens, TTFT, and “it just works” APIs; differentiation is custom kernels + capacity, not brand. Groq/Cerebras are hardware-tied; others are NVIDIA-heavy software shops.

C. Gateways, routers, and observability OpenRouter, LiteLLM, Portkey, Helicone, Martian. Dynamics: Thin but sticky control plane—failover, spend caps, A/B routing, evals. Winner is whoever becomes the default “inference mesh” for multi-model apps (less sure Martian’s long-term traction vs. OpenRouter/LiteLLM).

D. Specialized inference hardware Groq, Cerebras, SambaNova, NVIDIA (H100/H200/B200), AMD MI300. Dynamics: NVIDIA still takes most dollars; LPU/wafer-scale bets win only on extreme latency or batch-1. Software stack quality (not FLOPS) decides adoption.

E. Optimization / compilation / compression Neural Magic (Red Hat), NVIDIA TensorRT, GGML/llama.cpp ecosystem, Intel OpenVINO (edge). Dynamics: Quantization + speculative decoding + graph compilers are table stakes; value is shrinking as engines absorb these techniques. Neural Magic is the clearest independent “make it smaller/faster on commodity CPUs/GPUs” story.

(3) Who wins the most value. Serving engines + the inference clouds that productize them (A+B). Routing (C) takes a tax but is easier to swap; hardware (D) is cyclical and NVIDIA-dominated; pure optimization (E) gets absorbed. Durable value sits with whoever owns the request path (batching, KV cache, SLA) and the customer relationship for production traffic—Together/Fireworks-style platforms sitting on vLLM/SGLang/TRT-LLM, plus NVIDIA’s tax on every GPU token.

(4) White space (no incumbent owns this cleanly)

  1. Agent-native inference: serving that is first-class for long-running tool loops, speculative multi-step plans, and KV-cache sharing across tools/sessions—not just chat completions. Current engines optimize single-turn throughput.
  2. Heterogeneous SLA router with economic guarantees: a control plane that continuously places each request across NVIDIA / Groq / Cerebras / on-prem with contracted p99 latency and $ per token, including automatic fallback and true-up billing. Today’s routers are mostly API multiplexers, not a real-time market + SLA fabric.

Less sure: relative ranking of Martian vs. Portkey/OpenRouter; how much Cerebras/SambaNova convert design wins into recurring inference revenue vs. Groq.