SysadminNews

IBM Puts $240M Into a Together AI Inference Cluster

On this page
  1. The deal
  2. Why serving, not training
  3. The networking choice
  4. What runs on it
  5. IBM Cloud's actual position
  6. Sources and further reading

IBM and Together AI signed a $240 million multi-year agreement on August 11, 2026 to stand up a large inference cluster on IBM Cloud. The hardware is Nvidia HGX B300 systems, the Blackwell generation, wired together with Spectrum-X Ethernet, and reporting puts the order at around 2,000 B300 units with availability expected in the first quarter of 2027. What makes it worth a read is not the number. It is that the whole cluster is aimed at serving open weight models rather than training a frontier one, and that it lands on a cloud which is not one of the big three. Serving, not training, is where the money moved this week.

The short answer

IBM and Together AI signed a $240 million multi-year deal to build a large AI inference cluster on IBM Cloud. It runs on Nvidia HGX B300 systems from the Blackwell generation, connected with Spectrum-X Ethernet, with roughly 2,000 B300 units reported and availability expected in the first quarter of 2027. The cluster targets open weight model serving, including DeepSeek, MiniMax and Kimi, for enterprises trying to bring inference costs down.

$240Mmulti-year agreement, signed August 11, 2026
~2,000Nvidia B300 units reported for the cluster
Q1 2027expected availability
Answer card describing the IBM and Together AI agreement signed on August 11, 2026, a 240 million dollar multi-year deal for an inference cluster on IBM Cloud built from around 2000 Nvidia HGX B300 Blackwell systems connected with Spectrum-X Ethernet, aimed at serving open weight models and expected to be available in the first quarter of 2027.
The agreement in one card. Sources: Reuters and The Next Web, August 11 and 12, 2026. PNG

Most nine figure AI announcements this year have been about training capacity. This one is not, and the difference is the whole point.

The deal

IBM and Together AI signed a multi-year agreement worth $240 million, announced on August 11, 2026. IBM builds the cluster on IBM Cloud. Together AI runs inference on it.

The compute is Nvidia HGX B300 systems, carrying Blackwell generation accelerators. Reporting puts the order at around 2,000 B300 units. The fabric is Nvidia Spectrum-X Ethernet. Availability is expected in the first quarter of 2027.

Why serving, not training

Training is a capital event. You buy a cluster, you burn months of it, you get a model at the end and then the machines go looking for the next job.

Inference is different in a way that matters to anyone who budgets infrastructure. It recurs. Every request from every user costs compute, forever, for as long as the product has customers. Together AI is reported to handle something on the order of 400 trillion tokens a month, which is not a research workload, it is a utility.

That distinction explains the shape of this purchase. Nvidia positions Blackwell explicitly for inference, and money moving toward serving capacity is a signal that the demand is coming from applications with real users rather than from labs racing each other. The same current runs under the disaggregated inference work at AMD and Cerebras and under the half trillion dollar financing structures being assembled around the buildout.

The networking choice

Spectrum-X is Nvidia's Ethernet stack for AI traffic: its own switch silicon, SuperNIC endpoints, and congestion control tuned for the all to all patterns that collective operations produce.

Picking Ethernet over InfiniBand for an inference cluster is a reasonable call and worth understanding rather than skimming. Inference traffic does not look like large scale training traffic. The operational skill base for Ethernet is enormous compared to InfiniBand. And the monitoring, troubleshooting and automation that a network team already owns mostly transfers.

Checklist figure summarising the IBM and Together AI inference cluster: 240 million dollar multi-year agreement, Nvidia HGX B300 Blackwell systems, roughly 2000 units reported, Spectrum-X Ethernet fabric rather than InfiniBand, open weight models including DeepSeek MiniMax and Kimi, IBM Cloud rather than a top three hyperscaler, and first quarter 2027 availability.
What the cluster is made of, and what each choice signals. PNG

Networking is a modest share of what an AI cluster costs and a very large share of what goes wrong when one is deployed. A vendor picking the more widely understood fabric for a serving workload is a small decision on the invoice and a significant one in the operations room.

What runs on it

Open weight models. Together AI's platform serves DeepSeek, MiniMax and Kimi among others, and this capacity is explicitly for that category.

The commercial argument is straightforward. Enterprises that want their inference bill to stop growing, or want a clearer boundary around where sensitive data goes, or simply do not want their product roadmap tied to one lab's model releases, have been moving toward open weight serving. Two thousand Blackwell systems is a bet that this continues through 2027.

IBM Cloud's actual position

IBM is not going to out scale the top three clouds and does not appear to be trying. The pitch here is economics and data control: run open models more cheaply, with a clearer answer about where they run.

Coverage has also framed this against a European preference for infrastructure that is not locked to a single American provider, which is a real market signal even if IBM being American complicates the framing.

For anyone sizing inference capacity for next year, the useful conclusion is narrower and more concrete. Serving is turning into a market with more than three credible suppliers, and pricing pressure in that market shows up directly in what it costs you to run a model in production.

Sources and further reading

Frequently asked questions

What was actually signed?

A multi-year agreement worth $240 million, announced on August 11, 2026, under which IBM builds a large scale AI inference cluster on IBM Cloud for Together AI. The compute is Nvidia HGX B300 systems, which carry the newer Blackwell generation accelerators, and the fabric is Nvidia Spectrum-X Ethernet. Reporting puts the volume at roughly 2,000 B300 units, with availability expected in the first quarter of 2027. The dollar figure is not a vague partnership commitment, it is a purchase of serving capacity, which is why the timeline is stated in quarters rather than in press release adjectives.

Why is inference the interesting word in this deal?

Because it changes what the hardware is optimised for and who pays for it. Training is a capital event: you buy a cluster, you spend months, you get a model. Inference is an operating cost that recurs every single time a user does anything. Nvidia positions Blackwell explicitly for inference workloads, and Together AI is reported to process on the order of 400 trillion tokens a month, which is the shape of a serving business rather than a research one. When large sums start moving toward serving rather than training, it means the demand is coming from applications that already exist and have users.

What is Spectrum-X Ethernet doing here, and why not InfiniBand?

Spectrum-X is Nvidia's Ethernet stack tuned for AI traffic patterns, combining its switch silicon with SuperNIC endpoints and congestion control aimed at the all to all communication that collective operations generate. Choosing it over InfiniBand for an inference cluster is a defensible call: inference has different traffic characteristics from large scale training, Ethernet operations skills are far more common in the industry, and the tooling around it is familiar to network teams who have never run a fabric manager. Networking is a small share of the capital cost of an AI cluster and a disproportionate share of the deployment problems, so the choice matters more than the price suggests.

Which models will run on it?

Open weight ones. Together AI's platform serves models including DeepSeek, MiniMax and Kimi, and this cluster is explicitly for that category rather than for a proprietary frontier model. That is the core of the commercial argument. Enterprises looking to cut inference bills, keep sensitive data under a clearer boundary, or avoid a single vendor's model roadmap have been steadily moving toward open weight serving, and this deal is a capacity bet on that trend continuing through 2027.

What does this mean for IBM Cloud specifically?

It is a deliberate positioning move. IBM Cloud is not going to out scale the top three hyperscalers, so the pitch is economics and data control instead of raw footprint: cheaper serving of open models, with a clearer story about where the data sits. Coverage has also noted a European appetite for infrastructure that is not tied to a single American giant, which cuts in IBM's favour in some accounts and not in others. For anyone planning inference capacity in 2027, the practical takeaway is that serving is becoming a competitive market with more than three credible addresses.