Together AI — open-model cloud: inference, fine-tuning, GPU clusters
September 14 at 04:40 · $0.132 total
MEMORANDUM
To: Investment Committee From: [Venture Investor] Date: May 24, 2024 Subject: Investment Recommendation: Together AI (Hypothetical Series A)
---
Thesis
Together AI has the potential to be a fund-returner because it is positioning itself as the default infrastructure layer—the AWS—for the open-source generative AI ecosystem. The core thesis rests on a bifurcated AI future: while closed models (OpenAI, Anthropic) will capture the frontier, open-weight models (Meta’s Llama, Mistral) are rapidly closing the performance gap. Enterprises are increasingly demanding open models to avoid vendor lock-in, ensure data privacy, and drastically reduce inference costs.
However, deploying, fine-tuning, and scaling open models is a massive infrastructure headache. Together AI abstracts this complexity away. If open-source AI captures even 30-40% of enterprise workloads, the platform that serves as the primary routing, inference, and compute layer for these models will be a $10B+ company. Together AI is currently leading the race to become that platform.
Product & Wedge
Together AI provides a full-stack cloud platform for open models.
- The Wedge: Blazing fast, serverless inference APIs. Developers can ping Together’s API to access Llama 3, Mixtral, and dozens of other models instantly, paying a fraction of the cost of OpenAI. They have built proprietary routing and optimization software that makes their inference speeds some of the fastest in the industry.
- The Expansion: Once developers are hooked on cheap, fast inference, Together upsells them on fine-tuning services and, ultimately, dedicated GPU clusters. By offering a seamless pipeline from prototyping on serverless endpoints to deploying custom models on reserved compute, they capture the entire lifecycle of enterprise AI development.
Market & Competition
The market for AI infrastructure is massive but brutally competitive. Together AI sits at the intersection of model-as-a-service and raw compute provision.
- Direct Startups: Fireworks AI (fierce competitor on inference speed/cost), Anyscale (strong in Ray-based compute and open models), Replicate (stronger in image/video and hobbyist developers), and Groq (competing purely on inference speed via custom LPU hardware).
- The Hyperscalers: AWS (Bedrock), Azure, and GCP. These are the ultimate threats, as they already own the enterprise cloud relationships.
- Ecosystem Players: Hugging Face (the repository of models, which also offers inference endpoints).
Traction & Business Signal
- Developer Adoption: Strong public signal. They have tens of thousands of developers on the platform and are frequently cited as the default endpoint for open-source hackathons and early enterprise deployments.
- Partnerships: Backed by NVIDIA, which provides a critical strategic advantage in securing scarce GPU supply.
- Funding: (Publicly known) Raised over $100M in late 2023 led by Kleiner Perkins, followed rapidly by another $100M+ round led by Salesforce Ventures at a $1.25B valuation.
- Revenue & Margins: Unknown. While usage is visibly exploding, actual ARR, net retention, and gross margins are not publicly disclosed.
Risks
These are the three existential threats to the business:
- The Inference Race to the Bottom (Margin Collapse): Serving open-source models is inherently at risk of commoditization. If anyone can host Llama 3, the primary differentiators become speed and price. We are already seeing a price war between Together, Fireworks, and Anyscale. If Together cannot maintain a software-driven edge in compute efficiency, their gross margins will be crushed, turning them into a low-margin compute reseller rather than a high-margin software business.
- The Hyperscaler Distribution Advantage: AWS, Azure, and GCP are aggressively integrating open models into their native AI platforms (e.g., AWS Bedrock). Enterprises already have massive committed spend (EDPs) and security approvals with these giants. Together AI must prove it can win enterprise contracts against the frictionless default of AWS.
- The "Open Source Dries Up" Risk: Together AI’s entire business model relies on organizations like Meta, Mistral, and others continuing to spend billions training frontier open-weight models. If Meta decides the strategic value of open-sourcing Llama 4 or 5 is no longer worth the compute cost, Together AI is left with an aging inventory of models, and closed-source providers will monopolize the market.
Verdict: Invest
At a hypothetical Series A, Together AI is a clear Invest. They have the strongest team in the open-source infrastructure space, a distinct performance edge in inference optimization, and the critical GPU supply relationships required to scale. They are perfectly positioned to ride the tailwinds of enterprise open-source adoption.
The One Question I Need Answered First: What is your blended gross margin, and specifically, what are the unit economics of your serverless inference API versus your dedicated GPU clusters? (I need to know if the serverless API is a loss-leader burning cash to acquire users, or if their proprietary software stack actually allows them to generate venture-scale software margins on commoditized open models.)
Together AI investment memo (hypothetical Series A)
Thesis. Together could be a fund-returner if open-weight models become the default production stack for a large share of inference and customization spend. The company sits at the intersection of two durable shifts: (1) enterprises and startups want cheaper, more controllable alternatives to closed APIs, and (2) they still need someone else to run the messy GPU, serving, and fine-tuning layer. If Together becomes the default “open-model cloud”—reliable inference, fast fine-tuning, and dedicated clusters with a developer-friendly surface—it can capture a thick slice of a market that scales with token volume rather than seat licenses. Winner-take-most dynamics in inference (latency, price, reliability, model catalog) plus switching costs from fine-tunes and private clusters can produce a high-margin infrastructure franchise, not just another GPU reseller.
Product & wedge. Together sells inference (serverless and dedicated), fine-tuning, and GPU cluster access optimized for open models (Llama family, Mixtral-class, and a rotating catalog). The wedge is not “we have GPUs”; it is a full stack for people who have already chosen open weights: competitive tokens-per-dollar, low latency, OpenAI-compatible APIs, and a path from prototype → production → private cluster without rewriting the app. Fine-tuning and LoRA-style customization deepen lock-in. Dedicated clusters and reserved capacity address the customers who will not stay on shared serverless once spend and SLAs matter. The product is therefore a vertical cloud for the open-model workflow, not a horizontal IaaS clone.
Market & competition. The relevant market is AI inference + customization infrastructure, a subset of the much larger GPU cloud. Spend is growing with model capability and application volume; open weights expand the addressable set of buyers who refuse closed-model lock-in or need on-prem-like control. Real competitors include: hyperscalers (AWS Bedrock/SageMaker, Azure, GCP Vertex) who bundle closed and open models; specialized inference players (Fireworks, Groq, Cerebras inference, Inferless-class startups); GPU clouds (CoreWeave, Lambda, Crusoe, Voltage Park); and platforms adjacent to the workflow (Hugging Face Inference, Replicate, Anyscale, Modal, RunPod). Together’s differentiation is open-model depth plus a single vendor for inference + fine-tune + cluster, versus pure hardware speed (Groq) or raw GPU rental (CoreWeave) or closed-model gravity (OpenAI/Anthropic via hyperscalers). The risk is that “good enough” hyperscaler open-model serving plus price wars compresses the independent layer.
Traction & business signal. Public facts: Together launched in 2022–2023 around open-model inference; raised a 2023 seed/Series A-scale round (reported ~$20M) and subsequently much larger rounds (including a 2024 Series B reported at $102.5M and later capital well into the hundreds of millions). It has publicly named customers and partners in the developer and enterprise ecosystem and is widely used as an inference endpoint for open models. Exact ARR, gross margin, net retention, concentration, and GPU utilization are unknown from public sources. Growth is implied by successive large raises and product expansion into fine-tuning and clusters, but unit economics versus GPU depreciation and power are not disclosed. Competitive win rates versus Fireworks/Groq/hyperscalers are unknown.
Risks (the three that kill the deal).
- Commoditization of inference. If vLLM, TensorRT-LLM, and hyperscaler serving close the latency/price gap, Together becomes a thin API on rented GPUs. Gross margin collapses; customers multi-home. Fine-tuning helps only if it stays sticky and high-margin.
- GPU and capital intensity. Cluster businesses eat cash, face supply shocks, and can be undercut by CoreWeave-style specialists or hyperscaler reserved instances. A downturn in AI capex or a jump in HBM/power cost can destroy returns even with revenue growth. Together must prove it is software-margin, not a leveraged GPU trade.
- Open-model quality stall or closed-model lock-in. If frontier open weights lag closed models enough that serious production traffic stays on OpenAI/Anthropic/Google, Together’s TAM shrinks to cost-sensitive and research use. Alternatively, Meta/Mistral/others shipping “official” inference could disintermediate independent clouds. Either path makes Series A ownership of a mid-scale GPU cloud unattractive.
Verdict: pass at a hypothetical Series A, unless one question is answered first: What is current gross margin on inference (after GPU depreciation, power, and idle capacity) and what is net dollar retention on the top 20 customers? Without evidence that this is a software-like inference business rather than a utilization-sensitive GPU rental with a nicer API, the capital intensity and competitive set do not justify Series A risk for a fund-returner outcome. If margins and retention are already strong and expanding with fine-tunes/clusters, the thesis reopens. Until then, pass.