# Understanding AI in a Month 8: Why Scale Worked

Understanding AI in a Month — Day 8 · 2026-09-16

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

For a few years, the papers announcing the biggest AI models put the size right up front. May twenty twenty, GPT-three: a hundred and seventy-five billion parameters — the adjustable numbers inside the model. December twenty twenty-one, DeepMind's Gopher: up to two hundred and eighty billion.

Then in March twenty twenty-three OpenAI published the GPT-four technical report, and section two says this.

> this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar

> — *OpenAI, 'GPT-4 Technical Report', arXiv 2303.08774 (v1 15 March 2023; read in v6 of 4 March 2024), section 2 'Scope and Limitations of This Technical Report'*

The current documents are no different. OpenAI's safety page for GPT-six Astra, published on the third of September, has a section called Model Data and Training without a single quantity in it. Anthropic's system card for Claude Opus five gives a knowledge cutoff and no size. Google's model card for Gemini three Pro explains why total and active parameters are now two different numbers — and states neither. And on the tenth of September, DeepSeek published a model whose card gives five different parameter counts, depending on which thing you're counting.

So the number stopped being the headline. The part that gets missed: a year before OpenAI stopped printing it, a published experiment had already shown that the number, on its own, couldn't tell you which model was better.

Two claims get read into that story, and both are wrong.

The first: scaling laws are laws, like gravity, so machines improve on a schedule. The people who found the curves don't say that. The paper that made the phrase famous is Scaling Laws for Neural Language Models, by Jared Kaplan and colleagues at Johns Hopkins and OpenAI, January twenty twenty. It has an appendix headed Caveats, and it opens with this.

> At present we do not have a solid theoretical understanding for any of our proposed scaling laws

> — *Jared Kaplan, Sam McCandlish et al., 'Scaling Laws for Neural Language Models', arXiv 2001.08361v1 (23 January 2020), Appendix C 'Caveats', first bullet*

> As printed in the source: “At present we do not have a solid theoretical understanding for any of our proposed scaling laws.”

The second reading is the mirror image: the companies stopped publishing because the curves broke and scaling hit a wall. Keeping a number private tells you nothing about that. Researchers were still fitting these curves this summer, and the most careful test of them I know of is one OpenAI ran on its own model, in that same GPT-four report — the one that won't give you the size.

Yesterday was where the training text comes from. Today is the curve drawn over it.

Two ideas.

You have a method. You train it on some examples, then test it on examples you kept back, and count how often it's wrong. Then you train it on ten times as many and run the same test. Plot those points and you have a learning curve. A scaling law is nothing more exotic than that: a learning curve measured across a very wide range, and found regular enough to extend.

In two thousand and one, Michele Banko and Eric Brill at Microsoft Research measured one — partly, they wrote, because they wanted a better grammar checker. Their question was about budget. A better algorithm, or more text? The task was choosing between words people confuse, like then and than, where edited writing supplies the right answers for free. They collected a billion words, a thousand times the largest training set used on that problem before, and trained four standard methods on it.

> Note that the curves appear to be log-linear even out to one billion words

> — *Michele Banko and Eric Brill, 'Scaling to Very Very Large Corpora for Natural Language Disambiguation', ACL 2001 (aclanthology.org/P01-1005.pdf), section 3 'Learning Curve Experiments'*

> As printed in the source: “Note that the curves appear to be log-linear even out to one billion words.”

Log-linear is the word to get right. It doesn't mean the improvement speeds up. In their experiment, accuracy climbs by roughly the same step every time the text grows tenfold — so each next step needs ten times as much text. Steady gains, multiplying appetite.

Then the finding that made it a budget argument.

> At least for the problem of confusable disambiguation, none of the learners tested is close to asymptoting in performance at the training corpus size commonly employed by the field.

> — *same document, section 3, recorded with its opening hedge (the previous generation's record began at 'none of the learners' and dropped it)*

Nowhere near their ceiling, in other words.

Now the second idea, the one that moves money. Modern curves track loss: a score for how little probability a model gave to the word that actually came next, where lower is better. Loss tends to fall by a steady proportion each time the model or the data grows tenfold — inside the range somebody measured, and as long as the other one keeps up. And for a model that runs all of itself for every word, the arithmetic of training is roughly six times model size times tokens — so doubling both takes about four times the computing. With a fixed budget you have to choose — a bigger model, or a smaller one run over more text — before the run starts.

Kaplan's team measured that trade-off and answered: mostly buy model. In DeepMind's later summary, ten times the computing should buy a model five and a half times bigger and only one point eight times as much text.

In twenty twenty-two a DeepMind team led by Jordan Hoffmann trained more than four hundred models to check, and came back with a much more balanced recipe.

> for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled

> — *Jordan Hoffmann et al. (DeepMind), 'Training Compute-Optimal Large Language Models', arXiv 2203.15556v1 (29 March 2022), abstract*

> As printed in the source: “for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.”

Then they built it. Chinchilla: seventy billion parameters and one point four trillion tokens, on the same computing budget as Gopher, which had four times the parameters and three hundred billion tokens. The paper reports that Chinchilla beat Gopher and three other larger models.

That's the result that made it hard to rank two models by parameter count alone. Some training details changed too, but the headline change was how the budget was divided.

The thread runs longer than the boom. Two thousand and one, the grammar checker. Two thousand seventeen: a team at Baidu Research found the same kind of regular curve in translation, language modelling, images and speech, with exponents, in their words, yet to be explained by theoretical work. Then Kaplan, then Chinchilla.

Twenty twenty-three, the GPT-four report. OpenAI predicted GPT-four's final loss on an internal codebase — code OpenAI says wasn't in the training set — by fitting a curve to smaller models trained with up to ten thousand times less computing. They say they made the prediction shortly after the run started, and that it was highly accurate.

They also forecast a coding score on HumanEval, a public set of programming problems, with a separate curve. They set the hardest problems aside and sorted the rest into groups by difficulty. OpenAI says the forecasts came close for most groups — except the easiest, where GPT-four came in under what they'd predicted. The next sentence reads:

> Certain capabilities remain hard to predict.

> — *same document, section 3.2, the sentence immediately after the easiest-bucket miss (new paragraph), introducing the Inverse Scaling Prize and Hindsight Neglect*

That's OpenAI's account of its own run, with its own smaller models, and outsiders aren't given what they'd need to repeat it.

Two postures to set beside that.

The first came the year before Kaplan's paper: Rich Sutton's essay The Bitter Lesson, dated the thirteenth of March twenty nineteen.

> The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.

> — *Rich Sutton, 'The Bitter Lesson', incompleteideas.net/IncIdeas/BitterLesson.html, dated 13 March 2019 on the page itself, opening sentence*

My reading is that it's about choosing methods that keep improving as computing grows, not about buying parameters — and Chinchilla fits it. Sutton himself, in a September twenty twenty-five interview with Dwarkesh Patel, asked whether language models will reach the limits of the data and be overtaken by systems that learn from experience. Later in the same conversation, asked about a world with billions of AI researchers, he set his own lesson aside and said this.

> That’s an empirical observation about a particular period in history.

> — *Dwarkesh Patel, interview with Richard Sutton, dwarkesh.com/p/richard-sutton, 26 September 2025, transcript at 00:49:51 (answering whether the Bitter Lesson would still apply in a world of many AI researchers); read 14 September 2026*

The second posture is the audit. In twenty twenty-four, inside eleven weeks, three groups went back over the founding papers. Epoch AI reported that the published numbers from one of Chinchilla's three estimates didn't match the other two — and that when they redid it themselves, it did. Tim Pearce and Jinyeop Song traced much of the gap between Kaplan and Chinchilla to Kaplan counting a different set of parameters, at small scale. A group led by Tomer Porian pointed to three details — the last layer's cost, the warm-up, and optimizer tuning — and reported that fixing them brought the two into agreement.

Those studies traced much of the disagreement to counting, tuning and fitting — in a recipe that, by the Chinchilla paper's own account, many big models of the day had been trained to.

Meanwhile, OpenAI's model page lists no sizes at all. Its three GPT-five-point-six tiers list the same limits and the same features; the price is what tells them apart.

Two things you can check yourself.

One. That OpenAI safety page I mentioned at the top, for GPT-six Astra — find the section called Model Data and Training. Today there isn't a single quantity in it. If a number turns up there, that's the company itself, in its own document, and it needs nobody's interpretation.

Two. Article fifty-one of the European Union's AI Act presumes a general-purpose model has high-impact capabilities when the computing used to train it, counted in floating-point operations, passes ten to the twenty-fifth. That automatic trigger counts compute, not parameters, though parameters are among the things the Commission can weigh. This July the Union amended dozens of provisions of that Act and left Article fifty-one as it was, including the power to move that threshold. Watch whether it gets used.

The phrase to keep is empirical scaling law. Empirical means measured: in twenty twenty, the people who found these curves wrote that they had no solid theoretical understanding of them, and later theory explains only parts.

For the training problem we've been following, the law's job is splitting a budget between a bigger model and more text. On the public evidence, it's been a good guide to that. That budget is the training run alone; the computing spent later, answering your question, is a separate bill.

What those training-budget curves predict is loss. Whether lower loss buys the ability you actually wanted is a separate question. Some abilities can be forecast with curves of their own — and the GPT-four report, which shows that working, also says others remain hard to predict.

So when somebody tells you a system is enormous, the useful reply is a question: how was the budget split, and what did you measure afterwards?

Tomorrow: why a model that can only continue text will answer your question instead.

That was day eight. Thank you for listening.

---

## Sources (31)

- Corinna Cortes, L. D. Jackel, Sara A. Solla, Vladimir Vapnik and John S. Denker, *Learning Curves: Asymptotic Values and Rate of Convergence* — 1993
- Michele Banko and Eric Brill, *Scaling to Very Very Large Corpora for Natural Language Disambiguation* — 2001
- Alon Halevy, Peter Norvig and Fernando Pereira, *The Unreasonable Effectiveness of Data* — March/April 2009
- Joel Hestness et al. (Baidu Research), *Deep Learning Scaling is Predictable, Empirically* — v1, 1 December 2017
- Richard Sutton, *The Bitter Lesson* — 13 March 2019
- Jared Kaplan, Sam McCandlish et al., *Scaling Laws for Neural Language Models* — v1, 23 January 2020
- Tom B. Brown et al. (OpenAI), *Language Models are Few-Shot Learners* — v1, 28 May 2020
- Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee and Utkarsh Sharma, *Explaining Neural Scaling Laws* — v1, 12 February 2021
- Jack W. Rae et al. (DeepMind), *Scaling Language Models: Methods, Analysis & Insights from Training Gopher* — v1, 8 December 2021
- Jordan Hoffmann et al. (DeepMind), *Training Compute-Optimal Large Language Models* — v1, 29 March 2022
- Aakanksha Chowdhery et al. (Google), *PaLM: Scaling Language Modeling with Pathways* — v1, 5 April 2022
- Pablo Villalobos et al. (Epoch AI), *Will we run out of data? Limits of LLM scaling based on human-generated data*, and the accompanying Epoch page — v1 26 October 2022, v2 4 June 2024; page 6 June 2024
- OpenAI, *GPT-4 Technical Report* — v1, 15 March 2023
- Niklas Muennighoff et al., *Scaling Data-Constrained Language Models* — v1, 25 May 2023
- Tamay Besiroglu, Ege Erdil, Matthew Barnett and Josh You (Epoch AI), *Chinchilla Scaling: A replication attempt* — v1, 15 April 2024; v2 read
- Tim Pearce and Jinyeop Song, *Reconciling Kaplan and Chinchilla Scaling Laws* — v1, 12 June 2024
- Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt and Yair Carmon, *Resolving Discrepancies in Compute-Optimal Scaling of Language Models* — v1, 27 June 2024
- Meta, *The Llama 3 Herd of Models* — v1, 31 July 2024
- DeepSeek-AI, *DeepSeek-V3 Technical Report* and *DeepSeek-V4* — 27 December 2024; 26 April 2026
- OpenAI, gpt-oss, GPT-5 and GPT-6 Astra pages, Deployment Safety Hub — 5 August 2025; 7 August 2025; 3 September 2026
- European Commission, *Guidelines on the scope of the obligations for providers of general-purpose AI models*, C(2025) 7719 final — 19 November 2025
- Google DeepMind, Gemini 3 Pro model card — released November 2025, last updated May 2026
- Roberts et al., *Test-Time Scaling Makes Overtraining Compute-Optimal* — v1, 1 April 2026
- Lovelace et al., *Prescriptive Scaling Laws for Data Constrained Training*; Bryant and Liu, *Practical Scaling Laws* — 2 May 2026; 9 May 2026
- Liquid AI, LFM2.5-350M model card — read 13 September 2026
- Regulation (EU) 2026/1744 (Digital Omnibus on AI), and the consolidated text of Regulation (EU) 2024/1689 — 8 July 2026; OJ 24 July 2026
- Moonshot AI, Kimi K3 model card; Z.ai, GLM-5.3 repository — read 16 September 2026
- Richard Sutton, interview with Dwarkesh Patel — 26 September 2025
- Anthropic, Claude Opus 5 system card, and model comparison table — 24 July 2026; read 13 September 2026
- Epoch AI, Notable AI Models dataset, and Trends dashboard — updated 11 September 2026; updated 5 February 2026
- OpenAI, model documentation — read 13 September 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
