GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

L&T Deploys India's First 10,000-GPU Single NVLink Cluster for Together AI in Chennai

India’s largest engineering conglomerate, Larsen & Toubro, has secured a contract valued at roughly $1.05 billion to $1.57 billion to deploy 10,000 NVIDIA B300 “Blackwell Ultra” GPUs at Vyoma.AI’s Chennai data center campus. The customer is Together AI, the San Francisco-based open-source AI cloud provider that crossed $1.15 billion in bookings earlier this year.

The number is not the story. The architecture is.

Single-Cluster vs. Distributed: The Distinction That Matters

India has had large GPU deployments before. A distributed deployment of 10,000 GPUs spread across independently networked servers and a unified single-cluster of 10,000 GPUs connected by a shared NVLink fabric are not the same machine. The difference in capability is not incremental — it is categorical.

When NVIDIA ships the B300 in its DGX B300 server configuration, eight B300 SXM GPUs connect to each other through NVLink 5, NVIDIA’s fifth-generation interconnect, delivering 1.8 terabytes per second of bidirectional bandwidth per GPU. A NVSwitch chip inside each chassis provides full all-to-all connectivity within the node. In a true single-cluster deployment, that fabric extends across every node — any GPU can address any other GPU’s memory at full NVLink speed without going through the slower Ethernet or InfiniBand hop that separates typical distributed racks.

At 10,000 B300 GPUs, the full-cluster interconnect means the entire machine behaves as a unified memory space for training runs. Models that would otherwise require careful gradient synchronisation across distributed nodes can instead run as single large-memory training jobs. That is the compute architecture that produces frontier model training runs — not GPU counts spread across cabinets.

Why Together AI and Why Chennai

Together AI operates one of the largest open-source inference platforms globally, running models including Llama, Qwen, and DeepSeek variants for developers who do not want to route traffic through closed hyperscaler APIs. The company crossed $1.15 billion in contracted bookings in Q1 2026, driven partly by enterprise demand for non-OpenAI, non-Anthropic inference at scale.

Chennai is a deliberate choice. The city has historically been India’s hardware and manufacturing capital, with established power infrastructure and a deep engineering talent base. The Vyoma.AI campus provides the physical density required to house 10,000 GPUs as a single machine — air cooling and power density requirements for B300 deployments are substantially higher than prior GPU generations.

India’s AI Infrastructure Inflection

India’s sovereign AI ambitions have produced a series of GPU deployment announcements over the past 18 months, most of them distributed clusters at government-affiliated institutions. The L&T-Together AI deployment is the first in India to make the single-cluster claim explicitly and to tie it to a specific customer with publicly disclosed economics.

The contract value of $1.05B to $1.57B (approximately $105,000 to $157,000 per GPU at B300 pricing plus infrastructure) is consistent with Blackwell Ultra build costs at scale. L&T has not disclosed the deployment timeline, but Together AI’s existing capacity commitments suggest the cluster is intended to go live within 12 months of contract signing.

Key Numbers

  • GPUs: 10,000 NVIDIA B300 Blackwell Ultra
  • Contract value: $1.05B to $1.57B (L&T-Vyoma.AI)
  • Customer: Together AI
  • Location: Vyoma.AI Chennai data center campus
  • Interconnect: Single NVLink 5 fabric
  • NVLink bandwidth: 1.8 TB/s bidirectional per GPU