GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

IBM and Together AI Sign $240M Deal for Open-Source AI Inference on IBM Cloud NVIDIA HGX B300

IBM and Together AI have signed a multi-year, $240M agreement to deploy a dedicated open-source AI inference cluster on IBM Cloud. The hardware: NVIDIA HGX B300 systems connected via NVIDIA Spectrum-X Ethernet networking. NVIDIA says the B300 configuration delivers 30x the AI factory output of prior-generation hardware. Target availability is Q1 2027.

The partnership structure is clean. Together AI brings the model library and open-source inference platform. IBM brings the regulated cloud footprint — enterprise relationships, compliance certifications, and a customer base that cannot run workloads on commodity clouds. IBM Cloud hosts the cluster; Together AI operates the inference service on top of it.

The Open-Source Inference Bet

Together AI’s founding premise is that open-source models are essential for the future of enterprise AI. The IBM deal is the largest infrastructure commitment that thesis has attracted to date. Together AI raised $800M in a Series C last month — this deal converts part of that capitalisation into contracted capacity rather than company-owned hardware.

For IBM, it addresses a gap in its AI portfolio. IBM Cloud has the enterprise relationships and compliance credentials. It has not had a first-party open-model inference story at scale. The Together AI agreement gives it one, at frontier hardware specs, without IBM needing to build or maintain the model layer.

Why HGX B300

The NVIDIA HGX B300 is the current-generation AI inference platform, succeeding the H100-based HGX configurations. The 30x output improvement figure covers a combination of chip performance, NVLink bandwidth, and Spectrum-X network throughput — the full AI factory stack, not just GPU compute.

Spectrum-X is NVIDIA’s Ethernet-native AI networking platform, designed for AI workloads that require high-bandwidth east-west traffic between GPU nodes. The IBM deployment will be among the first large inference clusters to run HGX B300 at scale for open-source model serving.

Market Context

The deal lands as enterprise demand for open-source model inference has split from frontier-closed-model consumption. OpenRouter data shows open-source models now account for 65% of token volume on the platform, with Chinese models taking 30% of US enterprise token use. Together AI is positioned in the middle of that shift.

The IBM partnership gives it a path into regulated industries — financial services, healthcare, government — where enterprise buyers need SLAs, data residency guarantees, and audit trails that commodity inference APIs cannot provide. Q1 2027 is when that path goes live.