Investment memo / Battle
Investment memo · scheduled battle

Together AI — open-model cloud: inference, fine-tuning, GPU clusters

September 14 at 04:26 · $0.141 total

gpt-5.6-luna-pro wins
judged by claude-fable-5 · both position-swapped passes agree
“B is more factually accurate (founder, round sizes, valuation, Kleiner-led Series A), names a broader real competitor set, and clearly flags unknown metrics.”
deepseek-v4-pro
$0.00637s1896 tokens

MEMORANDUM

To: Investment Committee From: [Venture Investor] Date: October 26, 2023 Subject: Together AI – Series A Investment Consideration Confidential

Thesis: The Arbitrage Event in Open-Source Compute

The venture-scale return in AI has, until recently, been captured exclusively by proprietary model owners (OpenAI, Anthropic). However, the marginal cost of high-quality intelligence is collapsing due to the open-source movement (Llama, Mistral). The bottleneck is no longer the model weights; it is the infrastructure required to run them.

Together AI represents a bet on a "picks and shovels" business that monetizes the deployment layer of the open-source ecosystem. The fund-returner thesis rests on one premise: proprietary model APIs will face severe margin compression from open-source alternatives, and the majority of enterprise AI traffic will flow through neutral, performance-optimized cloud providers rather than hyperscaler walled gardens. If Together AI becomes the "Cloudflare for LLMs"—the default layer for serving and fine-tuning open models—it captures a disproportionate share of the rapidly expanding inference market, creating a durable, high-velocity revenue engine.

Product & Wedge: Inference-First, Infrastructure-Bottom

Together AI’s strategy is a bottom-up stack play. Their wedge is not a slick UX; it is raw engineering performance.

  1. The Wedge (Inference): They have optimized the serving of open-source models (Llama, Mistral, Yi, etc.) to run faster and cheaper than generic GPU clouds. By utilizing custom kernels (FlashAttention-3, custom CUDA graphs) and aggressive batching, they can serve a model like Llama-3 at a fraction of the cost of naive vLLM deployments. This attracts startups and developers who want the quality of GPT-4-ish performance without the margin stack of OpenAI.
  2. The Expansion (Fine-tuning): Once data flows through their inference engine, the natural next step is customization. Together offers fine-tuning (LoRA/full) on the same hardware. This increases stickiness; the fine-tuned model artifact lives in their environment.
  3. The Endgame (GPU Clusters): At the top of the stack, they offer "GPU Clusters" for enterprises that want to own the full stack but not manage the physical hardware or Kubernetes layer. This moves them from a variable revenue model (per token) to a high-value, contracted, reserved revenue model (per GPU-hour).

Market & Competition: The Battle for the Neutral Layer

The market is bifurcating into "Model Builders" (OpenAI, Anthropic, Google) and "Model Runners." Together AI aims to be the leading neutral runner.

The Competitive Arena:

  • Hyperscalers (AWS, Azure, GCP): They own the raw compute. However, they are optimized for general-purpose workloads, not the high-throughput, low-latency requirements of transformer inference. They also face a conflict of interest (pushing their own proprietary models).
  • Pure-Play Inference Startups:
  • Fireworks AI: The most direct competitor. Similar pitch (fast, cheap gen-AI inference). Fireworks has been aggressive on the developer experience and compound AI systems (function calling).
  • Replicate: Focuses heavily on the "run open-source models via API" niche but is generally considered less performant on massive scale enterprise workloads; more of a hobbyist/prosumer tool historically.
  • Anyscale (Ray): Owns the orchestration layer, but is moving up the stack to serve models.
  • Neoclouds (CoreWeave, Lambda): They rent raw GPUs. Together AI is effectively a software layer on top of these providers. The risk here is Neoclouds decide to build their own inference software (CoreWeave has already moved this direction with its object storage and serving layers).

Differentiation: Together’s legacy (founded by the creators of the FlashAttention algorithm and the RedPajama dataset) gives them a technical credibility that pure resellers lack.

Traction & Business Signal

Note: Specific revenue and churn metrics are private and unknown to the public markets.

  • Technical Adoption: Together AI consistently ranks at the top of the Artificial Analysis LLM performance leaderboards for open-model serving throughput and latency. This is a strong technical signal.
  • Model Count: They serve hundreds of open-source models, indicating breadth.
  • Funding Velocity: This round (Series A) follows quickly on the heels of a substantial seed led by Lux Capital and Kleiner Perkins, suggesting strong insider conviction and likely revenue inflection (standard growth-stage pattern).
  • Ecosystem: They are the default backend for many popular open-source frameworks and have strong marquee logos (specific logos listed on their site, though exact ARR per logo is unknown).

Risks: The Three Deal Killers

  1. The Hyperscaler Margin Squeeze: AWS/Azure/GCP view inference as a commodity. If they subsidize their own serving costs (or bundle credits) to prevent churn to Together, Together's margin profile collapses. Unlike Snowflake, which battled AWS on data warehousing, inference is a much newer and less sticky workload. If AWS Bedrock offers superior latency or enterprise discounting, the "neutral party" advantage may not be enough to justify a premium.
  2. The Fragmentation of the Model Layer: The open-source ecosystem is chaotic. A new "Llama-killer" drops every quarter. While Together benefits from churn (users constantly testing new models), it also prevents optimization. The edge they built for Llama-3 might be wasted if the market shifts to a new architecture (e.g., Mamba/State Space Models) that requires different kernel optimizations. Unlike a proprietary model company, they are entirely dependent on the quality of OSS releases they do not control.
  3. Capital Intensity vs. Software Multiples: To offer GPU Clusters, Together must sign massive leases with CoreWeave or hardware vendors. This is a capital-intensive balance sheet game. If the market prices Together as a "hardware reseller" (low multiple) rather than a "software platform" (high multiple), the return profile deteriorates. The "GPU Cluster" revenue might be huge, but if the gross margin on that revenue is 10-15% (typical pass-through), it will dilute the valuation multiple even if it pumps the top-line ARR.

Verdict: PASS (Conditional)

Decision: We should pass at the current Series A terms unless one critical question is answered satisfactorily.

Rationale: The market opportunity is enormous, and the team is technically elite. However, the combination of hardware passthrough revenue and intense pressure from hyperscalers creates a "squeezed middle" scenario. If Together is primarily a reseller, we are overpaying for a hardware business.

The One Question: “Can you show me the gross margin differential between your Inference API revenue and your GPU Cluster revenue, and prove that Inference revenue is scaling faster than the Cluster revenue as a percentage of the mix?”

If the answer is that Inference/Software margins are >75% and growing faster than the raw compute rental business, the investment thesis holds, and we should lean in heavily. If the mix is shifting toward GPU Clusters, we are simply funding a capital-intensive infrastructure arbitrage with software-level risk.

gpt-5.6-luna-pro · winner
$0.00726s4708 tokens

Thesis

Together AI is building the infrastructure layer for companies that want to use open-weight AI models without owning the underlying GPU, serving, and optimization stack. The investment case is not simply “GPU cloud plus API.” It is that open models are likely to become a large and heterogeneous ecosystem—requiring specialized inference, fine-tuning, model evaluation, and dedicated compute—and Together could become the default operating platform for that ecosystem.

This could be a fund-returner if three conditions hold. First, open models take meaningful share from proprietary APIs in production, particularly where customers care about cost, latency, data control, or model customization. Second, Together’s software materially improves utilization and performance relative to generic clouds. Third, the company expands from an API business into high-value, recurring infrastructure contracts and captures substantial compute spend. A winner could resemble a combination of Hugging Face, an inference platform, and a specialized AI cloud—albeit with much more capital intensity.

The opportunity is large, but the key question is whether Together owns differentiated infrastructure or is primarily reselling scarce GPUs behind a polished developer experience.

Product & wedge

Together offers hosted inference for open models, fine-tuning, model training, GPU clusters, and related tools for deploying and optimizing generative AI workloads. Its platform supports popular open models such as Llama and Mistral, as well as image and multimodal workloads. The company emphasizes optimized inference, dedicated endpoints, high-throughput serving, and access to GPU infrastructure.

Its initial wedge is credible: developers want to experiment with open models but do not want to provision GPUs, optimize kernels, manage distributed systems, or operate production endpoints. Together can offer a faster path from model selection to production than assembling infrastructure on AWS, Google Cloud, or Azure.

The more valuable wedge is enterprise customization. Fine-tuning and dedicated clusters can create larger contracts and increase switching costs. If Together’s software lets customers serve a model at materially lower cost or latency than hyperscaler infrastructure, that performance can support attractive gross margins even when the company leases or purchases expensive GPUs.

Market & competition

The market spans several overlapping categories: GPU cloud, model APIs, MLOps, and AI application infrastructure.

Direct and near-direct competitors include CoreWeave, Lambda, Crusoe, and RunPod in GPU infrastructure; Fireworks AI, Modal, Baseten, Replicate, and Anyscale in model serving and deployment; and Hugging Face in open-model distribution, inference, and fine-tuning. Databricks and Snowflake increasingly offer model customization and serving to their installed enterprise bases. The hyperscalers—AWS Bedrock and SageMaker, Google Vertex AI, and Microsoft Azure AI—are the most important competitors because they control procurement, identity, networking, and enterprise relationships.

Together’s differentiation is its open-model focus, broad model support, and integration of compute, fine-tuning, and inference. That focus can be an advantage while open-model demand is fragmented and fast-moving. It is also a vulnerability: the same models are increasingly available through every major cloud, while customers may prefer to consolidate infrastructure with their incumbent provider.

Traction & business signal

Publicly known signals are strong but incomplete. Together AI was founded in 2022 by Vipul Ved Prakash and has raised substantial institutional backing, including a reported $102.5 million Series A in 2023 led by Kleiner Perkins, with participation from investors including Nvidia and Emergence. In 2024, the company announced an additional $106 million financing at a reported valuation of approximately $1.25 billion, led by Salesforce Ventures.

The company has publicly announced partnerships and integrations across the open-model ecosystem and has attracted developer usage around hosted models, fine-tuning, and inference. It has also positioned itself as a provider of dedicated GPU clusters for larger customers.

Revenue, ARR, gross margin, net retention, customer concentration, utilization, backlog, and the proportion of usage that converts into contracted enterprise revenue are unknown from public information. The number of paying customers and the split between API consumption, dedicated infrastructure, and professional services are also unknown. Those metrics matter more than fundraising: high API volume can still produce poor economics if GPU utilization and pricing are unfavorable.

Risks

1. Commodity economics and hyperscaler compression. The business may be trapped between expensive GPUs and falling inference prices. If Nvidia supply improves, hyperscalers discount aggressively, or open-model inference becomes more efficient, Together could face declining prices without sufficient software differentiation. Capital intensity and GPU depreciation could turn growth into low-margin revenue.

2. Open-model demand may not become the dominant enterprise path. Proprietary models from OpenAI, Anthropic, and Google remain easier for many enterprises to adopt and often lead on quality, reliability, and tooling. Open models may win specific workloads without generating enough recurring infrastructure spend to support a venture-scale independent platform.

3. Lack of durable moat and customer concentration. Model hosting is relatively easy to compare and increasingly standardized. Customers can migrate workloads among Together, Fireworks, AWS, Azure, CoreWeave, and internal clusters. If Together’s advantage is mainly access to GPUs or temporary model availability, retention may be weak. A few large customers could also create material revenue and capacity concentration.

Verdict

Invest at a hypothetical Series A, but only conditionally. The market timing, open-model positioning, and potential for software-plus-compute expansion justify underwriting a high-upside position. However, I would not invest on developer adoption or fundraising momentum alone.

The one question I would need answered first is: What are gross margin, net revenue retention, and customer concentration for the production inference and dedicated-cluster business—and do those metrics improve as utilization scales? If Together can demonstrate durable enterprise retention and software-supported margins, it merits investment. If revenue is predominantly low-margin GPU resale or volatile API usage, I would pass.