Together AI — open-model cloud: inference, fine-tuning, GPU clusters
September 14 at 04:16 · $0.130 total
Investment Memo: Together AI
Company: Together AI — open-model cloud (inference, fine-tuning, GPU clusters) Stage under consideration: Hypothetical Series A Author: [Partner], [Fund]
---
Thesis
If open-weight models remain within striking distance of frontier closed models — and Llama, Mistral, Qwen, and DeepSeek suggest they will — then a large share of enterprise AI workloads will run on open models for cost, control, and data-privacy reasons. Those workloads need a serving layer. Together AI's bet is to be the "AWS of open models": a full-stack cloud spanning serverless inference, fine-tuning, and dedicated GPU clusters. The fund-returner case: inference is a compounding, usage-based revenue stream tied to the fastest-growing compute category in history; if Together captures even a few percent of open-model inference spend at scale, this is a multi-billion-dollar revenue company. Crucially, the team (Tri Dao, author of FlashAttention, is Chief Scientist; founders include Vipul Ved Prakash and Stanford's Chris Ré and Percy Liang) has genuine systems-research depth — the one durable edge in a business where speed and cost-per-token are the product.
Product & wedge
Three layers: (1) serverless inference APIs for 100+ open models with strong price/performance claims (FlashAttention, speculative decoding, custom kernels); (2) fine-tuning as a service, letting customers own their weights — a real differentiator versus OpenAI; (3) reserved GPU clusters (H100/H200/B200) for training. The wedge is developer-first inference: drop-in OpenAI-compatible endpoints at lower cost, converting API experiments into fine-tuning and cluster commitments. The research pedigree isn't marketing — kernel-level optimization directly determines gross margin and pricing power in a business where the product is literally tokens-per-dollar.
Market & competition
AI inference spend is plausibly a $50B+ market by late decade. But this is the most crowded lane in infrastructure:
- Direct open-model inference rivals: Fireworks AI, Baseten, Replicate, Anyscale, Groq (custom silicon), and hyper-aggressive price players like DeepInfra.
- GPU clouds moving up-stack: CoreWeave, Lambda, Nebius, Crusoe.
- Hyperscalers: AWS Bedrock, Azure AI, Google Vertex — distribution monsters bundling open models into existing enterprise contracts.
- Model providers themselves: Mistral's own API; Meta partnering directly with clouds.
- Existential adjacency: Nvidia (NIM microservices) could commoditize the serving layer.
Differentiation rests on research-driven performance and the full-stack (inference → fine-tune → cluster) motion. That's real but fragile: kernel optimizations diffuse (FlashAttention is open source; vLLM and SGLang are excellent and free), and the floor price of a token trends toward raw GPU cost.
Traction & business signal (public only)
- Raised a $102.5M Series A (Nov 2023, Kleiner Perkins), $106M (Salesforce Ventures, Mar 2024), a $305M Series B at ~$3.3B valuation (Feb 2025, General Catalyst/Prosperity7), and reports of a subsequent round at higher valuation.
- Reported (press) ~$100M+ annualized revenue run-rate as of early 2025 and hundreds of thousands of registered developers. Not audited; gross margin unknown — critical, since GPU resale revenue can carry thin or negative margins.
- Named customers/partners publicly referenced include Salesforce, Zoom, Washington Post, and various AI startups.
- Unknown: revenue mix between serverless inference (high multiple) vs. reserved cluster resale (low multiple), net revenue retention, customer concentration, GPU lease liabilities and their terms.
Risks — the three that kill the deal
- Margin structure is a mirage. If most revenue is GPU cluster resale, this is a capital-intensive middleman marked like a software company. CoreWeave trades on infrastructure multiples for a reason. A $3B+ valuation on $100M of ~20%-gross-margin revenue is not venture math. The serving-layer software premium only exists if inference is the majority of revenue and defensibly priced.
- Price war to zero. Open-model inference is nearly a commodity: identical weights, OpenAI-compatible APIs, switching cost measured in one config line. DeepInfra, hyperscalers, and eventually Nvidia can price at or below cost. Together's performance edge must stay ahead of open-source serving stacks forever to sustain margin. History of infrastructure commoditization says that's a hard treadmill.
- Open-model dependency. The entire business assumes frontier-competitive open weights keep shipping. If Meta deprioritizes open Llama releases (signals in 2025 were mixed) and Chinese models face enterprise procurement friction in the US, the quality gap versus closed models widens and workloads flow back to OpenAI/Anthropic. Together doesn't control its own supply of product.
Verdict
Invest at Series A — at Series A prices and terms, i.e., the 2023 ~$500M–$1B entry, not today's marks. The team is top-decile for exactly this problem, the market is enormous and early, and full-stack open-model serving is a legitimate shot at a generational infrastructure company. The bear cases are real but are execution/market-structure risks, not "this shouldn't exist" risks — acceptable at Series A pricing with 30–50x upside if inference margins hold.
The one question I need answered first: What is the revenue split and gross margin between serverless inference and GPU cluster resale — and what is net revenue retention on the inference cohort? That single disclosure tells you whether this is a software company riding compute growth or a leveraged GPU reseller wearing a software multiple.
Thesis
Together AI is a credible candidate for a fund-returning company because it sits at the intersection of three durable trends: the rapid adoption of open-weight models, the shortage and expense of specialized GPU capacity, and the migration of model development from research environments into production software. Its ambition is broader than being another GPU marketplace: Together provides inference, fine-tuning, training infrastructure, and managed GPU clusters in a single platform optimized for open models.
If it becomes the default operating layer for companies building on Llama, Mistral, Qwen, DeepSeek, and successor models, the upside is substantial. The company could capture infrastructure spend, model-optimization software revenue, and eventually high-value enterprise workloads. The key venture-scale possibility is that Together compounds proprietary advantages from workload scheduling, kernels, model serving, and customer usage data while becoming a trusted alternative to hyperscalers.
The opportunity is especially attractive if open models continue taking share from proprietary APIs. Customers want lower cost, more control, data residency, and the ability to customize models. Together is positioned to monetize all four needs.
Product & wedge
Together AI offers a cloud platform for open-source and open-weight AI. Its products include hosted inference APIs, fine-tuning, model training, and dedicated GPU clusters. It has also invested in optimized inference and training software, including high-performance serving for popular open models.
The initial wedge is developer access to high-quality open models without the operational burden of sourcing GPUs, configuring distributed systems, or optimizing inference. This is a useful entry point: developers can move from experimentation to fine-tuning and then to dedicated production capacity without changing vendors.
The more defensible product is the integrated stack. A customer may begin with an API call, progress to fine-tuning, and ultimately require a private cluster. Together can potentially retain that customer across the full lifecycle. Its focus on open models should also allow it to move quickly as new architectures emerge, rather than being tied to one proprietary model family.
Market & competition
The market is large but structurally competitive. Cloud infrastructure spending is enormous, and generative-AI infrastructure is growing rapidly. However, gross margins and differentiation vary significantly between commodity GPU rental and software-enabled inference.
Direct and adjacent competitors include CoreWeave, Lambda, Crusoe, and Fluidstack for GPU infrastructure; AWS, Google Cloud, Microsoft Azure, and Oracle for broad cloud platforms; and Fireworks AI, Baseten, Modal, Replicate, Anyscale, and Hugging Face for model hosting, inference, and developer tooling. Databricks and Snowflake also compete for enterprise model-development budgets.
Together’s differentiation is its open-model specialization and claimed full-stack coverage. The concern is that the largest clouds can subsidize GPU services, while focused competitors may offer better economics or a sharper developer experience. The company must prove that optimization and workflow integration create enough switching costs to avoid becoming a thin-margin reseller of scarce compute.
Traction & business signal
Publicly known: Together AI was founded in 2022 by Vipul Ved Prakash and has raised substantial venture financing, including a publicly reported Series A of approximately $102.5 million led by Kleiner Perkins, with participation from investors including NVIDIA and Emergence Capital. The company has publicly announced partnerships and integrations around prominent open models and has attracted developers and enterprise customers to its inference and training platform.
The company has also publicly discussed rapid usage and revenue growth. Exact audited revenue, gross margin, net revenue retention, customer concentration, GPU utilization, and contribution margin are unknown. The percentage of revenue from inference versus dedicated infrastructure is unknown. The durability of reported growth, particularly after GPU supply conditions normalize, is unknown.
The strongest business signal is strategic relevance: major model and infrastructure investors have backed the company, and customers increasingly need a production-grade alternative to self-hosting open models. The weakest signal is the lack of detailed public unit economics. In this category, impressive top-line growth can conceal expensive GPU leases, idle capacity, and low-margin pass-through revenue.
Risks
1. Infrastructure economics kill the deal. If Together owns or commits to too much capacity before demand is predictable, utilization shortfalls and GPU depreciation can destroy margins. If it owns too little, it becomes dependent on upstream suppliers and cannot guarantee availability. The company needs software-level gross margins or exceptional utilization; otherwise, hyperscalers and specialized GPU clouds can compete away profit.
2. Platform commoditization kills the moat. Open-model hosting is attractive precisely because models are portable. Customers can move between Together, AWS, Fireworks, Replicate, and self-hosted infrastructure. Kernel optimizations can be copied, and popular models may increasingly ship with strong native serving stacks. Together must demonstrate that its orchestration, performance, reliability, and developer workflow create meaningful switching costs.
3. Model and customer concentration kill the upside. Open-model popularity is volatile. A new architecture, licensing change, or breakthrough proprietary model could reduce demand for the models Together specializes in. Conversely, a few large customers may account for disproportionate revenue. Unknown customer concentration and retention make this a material underwriting gap.
Verdict
Invest at a hypothetical Series A, but only with disciplined financing and a strong preference for software-led economics rather than balance-sheet-heavy GPU ownership. The category is strategically important, the product wedge is credible, and the company can plausibly become the control plane for open-model production.
The one question I would need answered first is: What are gross margins and contribution margins by product—managed inference, fine-tuning, and dedicated clusters—after all GPU, networking, and energy costs, and how do those margins change at mature utilization? Without a compelling answer, Together is a fast-growing infrastructure reseller, not a fund-returning software platform.