# Understanding AI in a Month 14: What a GPU Is Doing

Understanding AI in a Month — Day 14 · 2026-09-24

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

Sometime between the first and the eleventh of this month, NVIDIA changed the product page for Rubin, its newest chip. Archived copies of the page show the edit. Its main figures for the arithmetic AI models use stayed the same. The figure for how fast it can read its own memory went down, from twenty-two terabytes a second to nineteen point two, and so did the figure for its direct links to the other chips in its rack.

In the same edit, the headline fifty petaflops gained a footnote, sparse specification, which NVIDIA's June datasheet already said. And a note calling those figures preliminary disappeared. NVIDIA's datasheet and its engineering blog still say twenty-two, and the page doesn't say why it changed. In February, analysts at SemiAnalysis said they understood memory suppliers were having trouble meeting NVIDIA's requirements. That's their reading, not NVIDIA's explanation.

So the numbers that moved were how fast the chip is fed, not how fast it computes. Today: what these chips actually compute, and why the feeding decides how fast a model answers you.

Two readings of a chip headline, and both mislead.

The first is the count. At its conference in March last year, NVIDIA revisited the name of its top chip, two slabs of silicon in one package. Tom's Hardware reported the point in quotation marks.

> Blackwell was named wrong.

> — *Jarred Walton, 'Nvidia announces Rubin GPUs in 2026, Rubin Ultra in 2027, Feynman also added to roadmap', Tom's Hardware, report of NVIDIA's GTC 2025 keynote, published 18 March 2025, the third of its introductory paragraphs; the line is carried in quotation marks as a point made in Jensen Huang's keynote; https://www.tomshardware.com/pc-components/gpus/nvidia-announces-rubin-gpus-in-2026-rubin-ultra-in-2027-feynam-after, read 24 September 2026*

By September NVIDIA's own press release was using the name that counts the slabs, the dies. By January it was counting packages again, and The Register's Tobias Mann reported that NVIDIA seemed to have changed its mind. So "ten thousand G P Us" depends on who's counting, and what.

The second reading is the headline speed. NVIDIA's page for the H100 lists almost four thousand trillion operations a second on eight-bit numbers, with a footnote: with sparsity. NVIDIA's Blackwell datasheet says the ordinary figure, called dense, is half the sparse one; for Rubin, the sparse fifty sits over a dense thirty-five. And even the dense figure isn't how fast a model answers you.

First, what the chip computes. On day four, each word-piece became a list of numbers, passed through layers of learned weights. Each layer's weights form a grid, and pushing a list of numbers through a grid is a matrix multiplication: multiply each number by a weight, add up the products, over and over. Jared Kaplan and colleagues at OpenAI, in twenty twenty, put that at about two operations per parameter for every token, and said where the two comes from.

> the factor of two comes from the multiply-accumulate operation used in matrix multiplication

> — *Jared Kaplan and colleagues (OpenAI), 'Scaling Laws for Neural Language Models', arXiv 2001.08361 v1, 23 January 2020, section 2.1*

So yesterday's model, G L M five point three, with about forty billion parameters active on each token, does roughly eighty billion operations for every token.

Second, why a graphics chip. Drawing a screen means doing much the same small piece of arithmetic for millions of points at once, and matrix multiplication has that shape too: each number in the answer can be worked out separately, at the same time. In June two thousand and nine, Rajat Raina, Anand Madhavan and Andrew Ng, at Stanford, reported training large neural networks on a graphics card up to seventy times faster than on a dual-core processor. And in twenty seventeen NVIDIA built the idea into the silicon: tensor cores, units whose one job is to multiply small grids of numbers, four by four, and add the result.

Third, the one to hold: when a model is answering people, the arithmetic often isn't the limit. The Stanford paper timed a multiplication of two grids, each a thousand by a thousand, on its card, about twenty milliseconds in all, and added

> the actual computation takes only zero point five percent of that time

> — *Rajat Raina, Anand Madhavan and Andrew Y. Ng, 'Large-scale Deep Unsupervised Learning using Graphics Processors', Proceedings of the 26th International Conference on Machine Learning (ICML 2009), Montreal, June 2009, section 2 'Computing with graphics processors'*

> As printed in the source: “the actual computation takes only 0.5% of that time”

The rest was copying numbers onto the card from the computer's main memory, a different road from the one inside today's chips; the paper says memory on the card itself was fast. What carries over is the shape: arithmetic was cheap, moving numbers wasn't.

It had an older root. In a short note published in March nineteen ninety-five, William Wulf and Sally McKee argued that processors were speeding up much faster than memory, so the time spent waiting on memory would keep growing. Of system performance, they wrote

> In fact, it will hit a wall.

> — *Wm. A. Wulf and Sally A. McKee, 'Hitting the Memory Wall: Implications of the Obvious', ACM SIGARCH Computer Architecture News 23(1), pp. 20-24, March 1995 (note dated December 1994), p. 20, last sentence on the page; read in the Internet Archive capture web/2005 of www.cs.virginia.edu/papers/Hitting_Memory_Wall-wulf94.pdf on 24 September 2026*

Their note was about how long memory takes to answer. The picture used today, about how much it delivers per second, came from Berkeley in October two thousand and eight: the roofline, drawn by Samuel Williams, Andrew Waterman and David Patterson.

> We believe that for the recent past and foreseeable future, off-chip memory bandwidth will often be the constraining resource

> — *Samuel Williams, Andrew Waterman and David Patterson, 'Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures', UC Berkeley technical report UCB/EECS-2008-134, 17 October 2008 (published in Communications of the ACM 52(4), April 2009), section 3 'The Roofline Model', opening sentence*

The roofline asks one question of any job: how many operations does it do for each byte it fetches from memory? A chip has a break-even ratio, its peak operations per second divided by the bytes per second its memory delivers. Below that ratio the chip waits on memory; above it, arithmetic is the limit. For the H100, the break-even is about three hundred operations per byte on sixteen-bit numbers, and about six hundred on eight-bit.

Now put a model on it. To write one token for one person, the chip reads every active weight from memory and uses each one for about two operations. Stored at eight bits, a weight is one byte, so that's two operations per byte, against a break-even in the hundreds. G L M five point three's forty billion active parameters are about forty gigabytes per token. One H100 reads its memory at three point three five terabytes a second. Divide, and the ceiling is about eighty-four tokens a second, while its eight-bit arithmetic could keep up with roughly three hundred times that.

Yesterday's model shows why this matters. It's sparse in a different sense from that footnote: only some of its parts run for each word-piece. It still has to store all seven hundred and fifty billion parameters, but for each word-piece it writes for one person, it reads only the forty billion that run. That saves no storage, and a great deal of reading. The whole model doesn't fit on one H100, and serving many people at once changes the sums. Those are the next two days.

Here are the last nine years on NVIDIA's own datasheets, and these peak figures come only from the vendor.

The V100 of twenty seventeen: about a hundred and twenty-five trillion sixteen-bit operations a second, and nine hundred gigabytes a second from memory. Rubin, on its June datasheet: four thousand trillion sixteen-bit operations a second, dense, and twenty-two terabytes a second. At the same precision, compute grew about thirty-two times and bandwidth about twenty-four, or twenty-one on the product page's new figure. Not so different.

Count Rubin's dense four-bit figure instead, and compute grew about two hundred and eighty times: about thirty-two from more and faster sixteen-bit arithmetic, now on two dies instead of one, and about nine from smaller numbers. Nearly all of compute's lead over memory comes from those smaller numbers. And the gap that matters for writing one answer was already there in twenty seventeen: about a hundred and forty operations per byte to keep the V100 busy on sixteen-bit numbers, against about one for a model stored that way.

NVIDIA is plain about where the limit is. In July, its engineering blog on Rubin said

> The decode, or generation, phase of inference is fundamentally memory subsystem bound.

> — *Eduardo Alvarez, Vishal Mehta and Farshad Ghodsian, 'Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI', NVIDIA Technical Blog, 21 July 2026, section 'Co-designing a memory subsystem for maximum power and compute efficiency', first sentence; https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/, read 24 September 2026*

It adds that what counts is how much of that bandwidth the software actually uses, not the peak on the spec sheet.

AMD's page for its coming MI455X claims six per cent more memory bandwidth than Rubin, and its own table sets that against twenty-two. Against nineteen point two it would be about twenty-one per cent.

Google's newest pair of chips points the same way. Of the two it described in April, the one for serving answers, T P U eight i, has less four-bit arithmetic than the training chip, and more memory, and more bandwidth.

The posture that survives is to ask for dense operations per second at a named precision and bytes per second of memory, together, then look for somebody measuring what you care about. Artificial Analysis, which neither builds chips nor sells the model, measures G L M five point three writing about fifty-nine tokens a second on Z dot A I's own service, and up to about four hundred and twenty across twenty-two hosts, on hardware its page doesn't name. The fastest is five times the one-H100 ceiling. That doesn't break the arithmetic: the host doesn't say what it runs on, and it could be a newer chip, several chips, smaller numbers, or more than one token per pass. And MLCommons, the benchmark consortium, listed Rubin, new to its results, as in preview on September the sixteenth.

One. Which Rubin bandwidth NVIDIA stands behind as volume shipments arrive: nineteen point two on the product page, or twenty-two on the datasheet and the blog. And whether anyone independent measures it.

Two. Whether Rubin leaves preview in the next MLPerf Inference round. MLCommons put its last two rounds six months apart, which points to next spring; no date is published.

Three. Whether AMD updates its comparison, so that six per cent more bandwidth than Rubin becomes about twenty-one.

The idea to keep is operations per byte fetched from memory. The roofline paper calls it operational intensity; you'll often hear the looser name, arithmetic intensity. A chip is a very fast calculator at the end of a road from a warehouse. The petaflops describe the calculator. For writing one answer for one person, the road sets the pace.

So when a chip headline arrives, ask three things. Operations per second, dense, at what precision? Bytes per second from memory? And how many operations does the job do for each byte it reads?

To read more: Williams, Waterman and Patterson, Roofline, in Communications of the A C M, April two thousand and nine.

---

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
