# Understanding AI in a Month 13: Capability, Compressed or Routed

Understanding AI in a Month — Day 13 · 2026-09-23

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

Last month, by its own release notes, a Chinese lab called Z dot A I released a model anyone can download, called G L M five point three. Because the weights are public, nobody has to take the lab's word about its size: Hugging Face, the site that hosts them, counts the parameters in the files. Seven hundred and fifty-three billion.

Now add up the model it replaces, G L M five point two. Same number — not close, the same, down to the last parameter. The configuration files agree on every setting that describes the model's design.

The lab's page explains it in half a sentence.

> every gain comes from post-training

> — *Z.ai, 'GLM-5.3' model card, the opening paragraph under its title, repository created 25 August 2026; https://huggingface.co/zai-org/GLM-5.3, read 22 September 2026*

So: same size, same design, and the same price — a dollar forty in, four dollars forty out, per million tokens. The lab claims a fifty per cent gain on an in-house coding test; its page gives no way to check it.

Yesterday was what extra thinking buys. Today, what quality costs when a system gets smaller, or sparser.

Two readings of a model that size, and both mislead.

The first is the one the number invites: that seven hundred and fifty-three billion makes it about six times as big as Mistral Medium three point five, at a hundred and twenty-eight billion. But only about forty billion of G L M's parameters work on each token — each word-piece it reads or writes. Mistral's model is dense: all hundred and twenty-eight billion run every time. Per token, the smaller-sounding model puts about three times as many parameters to work.

The second reading is the opposite mistake: sparse means cheap. The evaluation firm Artificial Analysis, which neither built nor sells it, files it in its standard summary line as

> amongst the leading models in intelligence, but particularly expensive when comparing to other open weight models of similar size

> — *Artificial Analysis, 'GLM-5.3 (max): Intelligence, Performance & Price Analysis', artificialanalysis.ai/models/glm-5-3, comparison summary, read 22 September 2026*

Cheap to compute and cheap to buy are different things, and the confusion lives in the gap.

Two old ideas get a large model's ability into something cheaper to run. One squeezes it. The other routes around it. Different origins, different mechanisms, often confused.

Squeezing first, and the recipe is older than the paper usually credited. In August two thousand and six, at a data-mining conference, Cristian Buciluă, Rich Caruana and Alexandru Niculescu-Mizil published Model Compression, which states the whole idea in one line.

> The main idea behind model compression is to use a fast and compact model to approximate the function learned by a slower, larger, but better performing model

> — *Cristian Bucilua, Rich Caruana and Alexandru Niculescu-Mizil, 'Model Compression', KDD '06, Philadelphia, 20-23 August 2006, section 2*

The procedure: have the big slow model label an enormous pile of examples, then train a small fast model on those labels. The small model isn't learning from the world. It's learning from the big model's answers.

The name came nine years later, from Geoffrey Hinton, Oriol Vinyals and Jeff Dean. They called it distillation, and they were careful about whose idea it had been.

> A version of this strategy has already been pioneered by Rich Caruana and his collaborators

> — *Geoffrey Hinton, Oriol Vinyals and Jeff Dean, 'Distilling the Knowledge in a Neural Network', arXiv 1503.02531 v1, 9 March 2015, introduction*

What they explained, and generalised, is the part people get wrong. You'd assume the student just copies the teacher's answers. In their version the teacher gives more than an answer: its probabilities for every alternative, the wrong ones included.

> The relative probabilities of incorrect answers tell us a lot about how the cumbersome model tends to generalize

> — *Geoffrey Hinton, Oriol Vinyals and Jeff Dean, 'Distilling the Knowledge in a Neural Network', arXiv 1503.02531 v1, 9 March 2015, introduction*

And their example:

> An image of a B M W, for example, may only have a very small chance of being mistaken for a garbage truck, but that mistake is still many times more probable than mistaking it for a carrot

> — *Geoffrey Hinton, Oriol Vinyals and Jeff Dean, 'Distilling the Knowledge in a Neural Network', arXiv 1503.02531 v1, 9 March 2015, introduction*

> As printed in the source: “An image of a BMW, for example, may only have a very small chance of being mistaken for a garbage truck, but that mistake is still many times more probable than mistaking it for a carrot”

A label saying car teaches the student one thing. The teacher's whole spread of confidence — overwhelmingly car, faintly truck, essentially never vegetable — teaches it how the teacher thinks the world is arranged. The pattern of its mistakes is part of the lesson.

Now the second idea, which starts earlier. In nineteen ninety-one, Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton described a mixture of experts: several small networks, plus a gate that decides which of them handles which examples. Using that to save computation — running only the experts the gate picks — was being proposed by around twenty thirteen, and in twenty seventeen Noam Shazeer and colleagues at Google built it at scale. Their abstract puts it like this.

> Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation

> — *Noam Shazeer and colleagues, 'Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer', arXiv 1701.06538 v1, 23 January 2017, abstract*

That sentence names the two quantities this episode turns on. Capacity: how much a model holds. Computation: how much work it does per token. In an ordinary network they're welded together, because adding a parameter means running it. Conditional computation unwelds them.

Here's the machinery, in the model I opened with. Most of its layers hold two hundred and fifty-six small expert networks each. In each of those layers, for every token, a gate picks eight; one more, a shared expert, skips the gate and runs every time. The other two hundred and forty-eight sit that token out. They're parts of one model, not separate models taking turns.

And here's what that doesn't buy you. Memory. To run at speed, all of its roughly seven hundred and fifty billion parameters have to sit on the hardware, because the gate could ask for any of them next. Routing saves arithmetic. It saves nothing on storage.

The number to watch in that second idea is a fraction: how much of a model runs on each token. Shazeer's own largest models already sent each example to just four of up to a hundred and thirty-one thousand experts. Later, open models started printing both numbers. January twenty twenty-four, Mixtral: forty-seven billion reachable, thirteen billion used — about a quarter. That December, DeepSeek version three: six hundred and seventy-one billion, thirty-seven billion per token — about a twentieth. This February, today's model family: forty billion of roughly seven hundred and fifty. Among this year's sparse open models that list both counts, the share runs from about three per cent to about eleven.

For models built this way, the headline count stopped predicting what a query costs to compute.

That February paper also holds a surprise. Its final training stage names its teachers like this.

> the final checkpoints from the preceding training stages serve as teacher models

> — *GLM-5 Team (Zhipu AI and Tsinghua University), 'GLM-5: from Vibe Coding to Agentic Engineering', arXiv 2602.15763 v2, 24 February 2026, section 3.5*

The teachers are the model's own earlier checkpoints, there to recover skills that later rounds of training can wear down. Same teacher-and-student idea as two thousand and six — except the student does the writing, the teacher grades it, and nothing gets smaller.

In late August, the same lab released a second model that answers the question the other way. G L M five point three Flash: three hundred and twenty billion total, eighteen billion active. Its page says it starts from a newly trained base, with a redesigned architecture; it doesn't say what, if anything, it learned from the big one. On Artificial Analysis's index it scores three points lower, and running it through that index cost about a ninth as much.

On the same index, OpenAI's G P T six Sol, released on the twenty-second of September, scores forty-four point one at its second-highest setting — a little below G L M five point three's forty-four point eight at its highest. And Sol's prices are higher overall: two dollars in and ten out, against a dollar forty and four forty.

Yet the evaluator's cost per task was fifty-three cents for Sol against two dollars for the open model — about a quarter as much. Over the whole index, the open model wrote five times as many tokens as Sol, and read four times as many. The price list favoured the open model; the volume decided the bill.

That's what a given level of quality costs on this benchmark, and it is not the price per token.

Three things, each on a page you can open.

One. This lab has put out a new release roughly every two months since February. On the next one's page, check the parameter total, which the hosting site computes from the files, and the configuration file beside it. If both match again, that fits an unchanged base for a third release running — though matching files can't prove it.

Two. The active parameters field on Artificial Analysis's model pages. The large open models fill it in; for OpenAI's and Anthropic's newest models, it's blank. Watch whether either company ever publishes that number.

Three. Cost per task on the evaluator's page for this lab's next model, beside its count of output tokens. Watch whether the next model closes the gap in tokens, or only in price.

The idea to keep is conditional computation: a system in which not every part runs on every input. It's now the design behind the largest open models.

The habit that comes with it is asking for two numbers where you're handed one. Total parameters tells you how much a model holds, and with its precision, what it takes to store. Active parameters is a rough guide to the arithmetic it does for each token. For most of this field's history those moved together. They don't any more, and a figure that gives you only the larger one has described a warehouse and said nothing about a delivery.

Neither number is the price of an answer. That also depends on how many tokens the answer takes, and what the service charges. So when you next read that a model is enormous, ask how many of its parameters ran, how many tokens it used, and what the answer cost.

To read more, the encyclopedia has articles on mixture-of-experts, on DeepSeekMoE, and on DeepSeek version three.

Tomorrow: what a graphics chip is actually computing.

That was day thirteen. Thank you for listening.

---

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
