Together AI Raises $800M at $8.3B as Aramco and NVIDIA Bet on Open-Source Inference
Together AI raised $800 million in a Series C on July 1, valuing the open-source AI neocloud at $8.3 billion. That is 2.5x its Series B valuation from 16 months ago and makes it one of the most heavily capitalised pure-play inference platforms outside the frontier labs.
The round was led by Aramco Ventures — Saudi Arabia’s energy company venture arm — with NVIDIA, Vista Equity Partners, General Catalyst, Emergence Capital, Salesforce Ventures, Pegatron, SentinelOne’s S Ventures, March Capital, DTCP Growth, and Lux Capital participating. Alongside the equity raise, investors have committed a further 500 megawatts of compute capacity to be capitalised independently and earmarked for Together AI’s expected demand growth.
What the Numbers Say
Together AI reported $1.15 billion in annual bookings for the most recent quarter. That puts it well past the growth-stage inflection point: the company is no longer proving product-market fit, it is scaling a confirmed demand curve.
The funding timeline tells the same story compressed:
- 2023 Series A: $102.5M (Kleiner Perkins led, NVIDIA joined)
- February 2025 Series B: $305M at $3.3B
- July 2026 Series C: $800M at $8.3B
Each round roughly tripled the valuation. The accelerant is inference demand from enterprises that want to run open models at production scale without managing GPU infrastructure themselves.
Why Aramco
Aramco Ventures leading an AI inference round is not instinctive. The firm’s typical mandate is energy and industrial technology. The 500MW compute commitment alongside the equity signals the actual thesis: sovereign wealth from energy-producing states is treating AI compute infrastructure the way they treated LNG terminals and telecoms networks in the 1990s and 2000s — long-duration physical assets that generate recurring yield.
NVIDIA’s participation is structurally different. Together AI runs workloads on NVIDIA hardware. The more informative read is what NVIDIA is betting on: inference is recurring revenue where training is one-time spend. Backing the open-source inference layer means backing continued GPU utilisation after frontier labs finish their training runs.
What Together AI Is Building
The immediate expansion target is managed inference: enterprises call open-source models through Together AI’s API without provisioning hardware. Models in the catalog include Llama 4, Qwen 3.x, Mistral, DeepSeek V4, and dozens of fine-tuned variants.
On the technical side, the company has shipped FlashAttention-4 optimised for NVIDIA Blackwell, Together Megakernel, and together.compile — a compilation layer that brings kernel-level optimisation to production workloads. The pitch to engineering teams is six-times cost reduction versus self-hosted equivalents.
The 500MW compute commitment from investors — capitalised independently, not as part of the equity raise — suggests Together AI is planning for 50x capacity growth over five years, a number cited in previous company communications.
The Market Context
Open-source models have closed to within 85-90% of frontier proprietary performance on standard enterprise tasks. For most production workloads — content generation, classification, retrieval augmentation, document processing — the gap does not justify the cost difference. Together AI’s bet is that the price compression unlocks a substantially larger enterprise market than the one frontier API providers can address.
The investor composition confirms the thesis is maturing: sovereign infrastructure capital and GPU manufacturers do not write growth-stage checks into speculative plays. The 500MW of committed compute alongside the equity makes it a capital-intensive infrastructure play, not a software startup.