# Understanding AI in a Month 15: From One Chip to a Cluster

Understanding AI in a Month — Day 15 · 2026-09-25

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

On Wednesday, the research firm SemiAnalysis published the third edition of ClusterMAX, its rating of companies that rent out AI chips. It covers seventy-seven. The clusters it gets its hands on go through three phases of tests: set-up, performance and reliability. Of its tests of the chips' own multiplying, it says

> these microbenchmarks very rarely distinguish providers

> — *SemiAnalysis (Jordan Nanos, Sam Harshe, Samuel Kruse and colleagues), 'ClusterMAX 3.0', SemiAnalysis newsletter, published 23 September 2026, section 'Phase 2: Performance', under the heading 'GPU Compute'; https://newsletter.semianalysis.com/p/clustermax-30-the-industry-standard, read 25 September 2026*

What its write-up dwells on is what surrounds the chips: storage, the network, support, pricing, security, and reliability, which it calls the most important criterion for many of the biggest customers. On NVIDIA's big racks it runs some tests with the chips' fastest links switched off, to measure the network beyond them on its own. SemiAnalysis sells research to the industry it rates, and its tiers are its own assessment. But notice the question it's asking. Not how many chips a provider has, but how well it runs them.

Today: why thousands of chips, each very fast, spend so much of their time waiting for each other.

Two readings of a big cluster number, and both mislead.

The first is that sixteen thousand chips do sixteen thousand times the work of one. Meta's paper on Llama 3, first posted in July twenty twenty-four, reports its chips doing useful arithmetic at forty-three per cent of their peak on eight thousand of them, and forty-one per cent on sixteen thousand. It traces the drop to how the work had to be divided. Doubling the chips made each one a little less productive, with nothing broken.

The second is that the chips in a rack become one machine. NVIDIA's page for its GB200 rack says its seventy-two chips form a domain that

> acts as a single, massive G P U

> — *NVIDIA, 'NVIDIA GB200 NVL72' product page, section 'Unlocking Real-Time Trillion-Parameter Models', opening paragraph; https://www.nvidia.com/en-us/data-center/gb200-nvl72/, read 25 September 2026*

> As printed in the source: “acts as a single, massive GPU”

That's the vendor's description. In its successor, the GB300 rack, each chip's links to the others are rated at about one point eight terabytes a second, counting both directions. Its connection to the network beyond the rack is rated at eight hundred gigabits a second. If that counts one direction, as network cards are usually rated, it's two hundred gigabytes a second both ways: roughly a ninth. Anything that has to cross from one rack to another takes the slower road. And both roads are slower than yesterday's road, from a chip to its own memory: on the same racks, about eight terabytes a second.

There are two reasons to use more than one chip. The first is room: yesterday's model needs about seven hundred and fifty gigabytes just to hold its weights, more than any one chip has. The second is time. Meta put the training compute for its largest Llama 3 model at three point eight times ten to the twenty-five operations. One H100, running flat out at the sixteen-bit peak from yesterday, would need more than twelve hundred years. At the pace Meta's chips actually managed, about three thousand.

So the work gets split, and there are three ways to split it.

The simplest is data parallelism. Every chip holds a whole copy of the model and reads a different slice of the examples. Jeff Dean and colleagues at Google described training with many copies of one model in twenty twelve, though in the first of their two methods the copies didn't wait for each other. In today's large runs, the copies pool what they've learned after every step, and nobody starts the next step until that's done.

The second is tensor parallelism. Google split Transformer layers across chips in twenty eighteen, and in twenty nineteen a team at NVIDIA set out a simple version for language models, in a paper called Megatron-LM. Each layer's grid of weights is cut into pieces on different chips, so each does part of every multiplication. The chips have to combine their pieces within every layer, so in practice this stays inside the fastest links.

The third is pipeline parallelism, an assembly line. The first layers live on one group of chips, the next layers on another, and work passes down the line. Google's GPipe paper, first posted in November twenty eighteen, names the catch.

> partitioning introduces some idle time per accelerator, which we refer to as the bubble overhead.

> — *Yanping Huang and colleagues (Google), 'GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism', arXiv 1811.06965 v5, 25 July 2019 (first posted 16 November 2018), section 2.3 'Performance Optimization'*

When a batch starts, the later stations wait for work to reach them, and at the end the early ones sit idle. Cutting each batch into many small pieces shrinks the bubble; GPipe found it could then be ignored, though on paper it never quite reaches zero.

In twenty twenty-one, researchers from NVIDIA, Stanford and Microsoft set out how to use all three together: split each layer across the chips inside one server, run the pipeline between servers, and copy the whole arrangement to use the rest. Llama 3 split its layers eight ways inside each eight-chip server, the same pattern.

Now, why they wait. Dean's paper had already named the usual cause of disappointing speed-ups when one model is split across many machines.

> many machines waiting for the single slowest machine to finish a given phase of computation

> — *Jeffrey Dean and colleagues (Google), 'Large Scale Distributed Deep Networks', Advances in Neural Information Processing Systems 25 (NIPS 2012), section 3 'Model parallelism'*

In a synchronous step the slowest participant sets the pace for everyone, and one that fails can stop them all. Meta's paper says it plainly.

> the synchronous nature of training makes it less fault-tolerant—a single G P U failure may require a restart of the entire job.

> — *Llama Team, AI @ Meta, 'The Llama 3 Herd of Models', arXiv 2407.21783 v3, 23 November 2024 (first posted 31 July 2024), section 3.3.4 'Reliability and Operational Challenges'*

> As printed in the source: “the synchronous nature of training makes it less fault-tolerant—a single GPU failure may require a restart of the entire job.”

Over one fifty-four-day stretch of training its largest model, Meta counted four hundred and sixty-six interruptions. Four hundred and nineteen were unexpected, about one every three hours, and it traced about seventy-eight per cent of those to hardware, confirmed or suspected. For Llama 3, it says, more than ninety per cent of the elapsed time went to useful training. And a chip doesn't have to fail to hold things up.

> Even a single straggler can slow down thousands of other G P Us, often appearing as functioning but slow communications.

> — *Llama Team, AI @ Meta, 'The Llama 3 Herd of Models', arXiv 2407.21783 v3, 23 November 2024 (first posted 31 July 2024), section 3.3.4 'Reliability and Operational Challenges'*

> As printed in the source: “Even a single straggler can slow down thousands of other GPUs, often appearing as functioning but slow communications.”

So the useful measure isn't how many chips, but how much of their arithmetic becomes model. In twenty twenty-two, Google's PaLM paper gave that a name, model FLOPs utilisation: the useful work a run achieves as a share of what its chips could do at peak. PaLM reported forty-six per cent, on six thousand one hundred and forty-four of Google's TPU chips. Llama 3 reported thirty-eight to forty-three. That isn't the same as fifty-odd per cent wasted: the measure leaves out work the chips really do, like recomputing results to save memory. And both figures come only from the labs that ran them.

Here's how the hardware grew around that idea, mostly on the vendors' own figures.

In twenty nineteen, Megatron's servers joined sixteen chips at three hundred gigabytes a second, with a hundred out of the whole server. By the H100, a server's eight chips each had nine hundred gigabytes a second of fast links, and Meta gave each chip four hundred gigabits a second, a hundred gigabytes counting both directions, to the rest of its cluster.

So the fast domain went from sixteen chips, to eight in an H100 server, to seventy-two in NVIDIA's newer racks. But per chip, the fast links stayed about nine times the network connection, if the network ratings count one direction: on Meta's H100 cluster, on NVIDIA's GB300 racks, and on Rubin as NVIDIA first described it. NVIDIA's revised Rubin figures, fast links down and network up, put it nearer seven to one.

And a second NVIDIA page has just followed. Sometime between the fourteenth of this month and today, NVIDIA's NVLink page cut Rubin's figure per chip from three point six terabytes a second to three, and two times the previous generation became one point seven. Its engineering blog, last modified on the twenty-first of August, still says three point six, and doubling. The page doesn't say why.

Two other postures.

Google builds a much bigger fast domain, of a different shape. Its documentation for its Ironwood chips describes pods of nine thousand two hundred and sixteen chips, each wired to its neighbours rather than to all the others, against seventy-two in NVIDIA's rack.

The other posture is to work around a slow road. In December twenty twenty-four, DeepSeek reported training its V3 model on two thousand and forty-eight of NVIDIA's H800 chips, whose fast links, by DeepSeek's own figures, were only about three times the cluster's network. It wrote its own pipeline schedule, DualPipe, so the chips had work to do while data was moving.

In June, the benchmark consortium MLCommons made that same model the largest in its training benchmark. CoreWeave said its time to hit that benchmark's target, a test rather than a full training run, fell from five point five four minutes on two thousand and forty-eight chips to two point oh two on eight thousand one hundred and ninety-two, and called that near-linear. Four times the chips bought about two point seven times the speed, and CoreWeave's statement doesn't say where the rest went.

One. Which NVLink figure NVIDIA stands behind for Rubin: three terabytes a second on its product pages, or three point six on its blog. And what the first independent tests of its racks measure.

Two. The next MLPerf Training round. MLCommons's announcement puts the next submission round in October and gives no date for results. When they do, watch whether spreading one job over more than eight thousand chips keeps paying.

Three. AMD's rack design, Helios, seventy-two of its chips in one domain, which other companies build. AMD says volume deployments are expected in the second half of this year; watch whether it appears in a public benchmark or rating by the end of December.

The idea to keep is utilisation: the share of a cluster's peak arithmetic that becomes model. A cluster is a team that moves at the pace of its slowest member, over roads of two speeds.

So when a lab announces a number of chips, ask three things. How big is the fast domain? How fast is the road between domains? And what share of peak did the run actually use?

To read more: Deepak Narayanan and colleagues, Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM, twenty twenty-one.

---

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
