# Understanding AI in a Month 3: One Token at a Time

Understanding AI in a Month — Day 3 · 2026-09-06

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it.*

---

Watch one of these systems answer. The words arrive in pieces, left to right, at about the pace of somebody typing. It is tempting to read that as an interface flourish — a progress bar with better manners.

It is more than that, and less than a window onto the machine's thinking. Run the ordinary way, the thing does build its answer one piece at a time, each piece chosen with the earlier ones fixed. But there is a second half to that sentence, and almost every popular explanation drops it. It writes one piece at a time. It does not *read* one piece at a time — it reads everything in front of it, at once, every step. This morning I measured how much that second half is doing.

Two readings, both wrong, failing in opposite directions.

The first is deflationary. *It's just autocomplete.* Gary Marcus put a version of it into WIRED in December twenty twenty-two: large language models, he wrote, are little more than autocomplete on steroids. Something in that is real — the objective genuinely is predict the next piece. What is wrong is the scope. Zhicheng Lin put it exactly last month: accurate, and misleading the moment it is taken for the whole account.

So I put a number on the gap. I gave two models — a small one from twenty nineteen, a four-billion-parameter one from this year — a two-token window: the previous two pieces of text and nothing else, roughly what the word autocomplete conjures. Then I compared each one's next word against what the same model says with the whole page in front of it.

They disagreed with themselves about four times out of five: twenty-three per cent agreement for the older model, twenty-two for the newer, on a passage neither could have seen. Cut the window to two and you have a different predictor.

The second reading runs the other way, and the moving cursor encourages it: the words appear one at a time because the thing is thinking them one at a time, in front of you. The display does not show you the boundaries of its deliberation. It is a delivery decision.

Yesterday we established what arrives at the model: a run of symbols from a frozen list. Today is what it does with them.

Two ideas. The first is **conditional probability**: the odds of the next thing depend on what came just before.

Its root is older than computers. In January nineteen thirteen the Russian mathematician Andrei Markov lectured in St Petersburg about counting he had done by hand: twenty thousand letters of Pushkin's *Eugene Onegin*, each classified vowel or consonant, in pairs. A letter is a vowel about forty-three per cent of the time; after a vowel, about thirteen; after a consonant, about sixty-six. His own summary:

> As we can see, the probability of a letter being a vowel changes considerably depending upon which letter – vowel or consonant – precedes it.

> — *A. A. Markov, 'An Example of Statistical Investigation of the Text Eugene Onegin Concerning the Connection of Samples in Chains', lecture of 23 January 1913, in the English translation by Gloria Custance and David Link, Science in Context 19(4), 2006, pp. 591-600, doi:10.1017/S0269889706001074, p. 596*

One correction, because the popular telling gets this wrong: Markov built no language model and was not studying language. His paper argues about a dispersion coefficient; Pushkin was test material.

The second idea is what happens when you run a measurement like that *forwards*.

Claude Shannon did that in nineteen forty-eight, in the paper that founded information theory. Section three is the same ladder, climbed — from letters drawn at random, up through each letter chosen using the one before it, to each word chosen using the word before it. His verdict:

> The resemblance to ordinary English text increases quite noticeably at each of the above steps.

> — *Claude E. Shannon, 'A Mathematical Theory of Communication', Bell System Technical Journal 27, July and October 1948, section 3 'The Series of Approximations to English', p. 7 of the reprint with corrections*

And then the sentence that answers today's real question:

> Note that these samples have reasonably good structure out to about twice the range that is taken into account in their construction.

> — *Claude E. Shannon, 'A Mathematical Theory of Communication', Bell System Technical Journal 27, July and October 1948, section 3 'The Series of Approximations to English', p. 7 of the reprint with corrections*

Coherence outruns the window: a two-word memory produced a ten-word run Shannon called not at all unreasonable. And he had no computer. For the higher rungs he opened a book at random, picked a letter, opened to another page, read until that letter turned up, and recorded whichever followed. He stopped for a stated reason.

> It would be interesting if further approximations could be constructed, but the labor involved becomes enormous at the next stage.

> — *Claude E. Shannon, 'A Mathematical Theory of Communication', Bell System Technical Journal 27, July and October 1948, section 3 'The Series of Approximations to English', its closing paragraph, p. 8 of the reprint with corrections*

Here is the machine, and it is smaller than people expect. At each step it scores every entry in its vocabulary — in one model I ran this morning, two hundred and forty-eight thousand of them, together. Those become probabilities, something picks one, the pick is appended, and the whole thing runs again. The model's own output becomes its input. The field's word for that loop is **autoregression** — though Jurafsky and Martin footnote their own vocabulary there: strictly the term means something narrower and linear, and language models are not.

The labour is no longer enormous, so this morning I climbed the rest of Shannon's ladder — same models, same passage, changing only how far back each could look: one piece, two, four, on up to two hundred and fifty-six. It falls the whole way, every rung, both models. The more of the past it sees, the less surprised it is.

So: does a chatbot arrive word by word because the machine works word by word? Half — and the half people get wrong is the checkable half.

Writing is sequential in the ordinary loop: step four's input contains step three's output. I counted it to be sure — run the plain way, twenty new pieces cost exactly twenty passes through the model, one each.

Reading is not sequential at all. I handed the same model a twenty-four-piece sentence in a single pass and asked what came back: twenty-four predictions, one for every position, computed together. Jurafsky and Martin say this plainly in their chapter on training — every position is scored at once against its true next piece, so one pass yields as many training examples as there are positions. That asymmetry is why these things could be trained at all: sequential reading would have made a trillion words of training a trillion passes.

Showing you the writing as it happens is a third thing, separate and optional. OpenAI's own documentation, on what happens if you do nothing:

> By default, when you make a request to the OpenAI API, we generate the model's entire output before sending it back in a single HTTP response.

> — *OpenAI, 'Streaming API responses', OpenAI API documentation, opening paragraph; https://developers.openai.com/api/docs/guides/streaming-responses, read 6 September 2026*

Show people a process and it acquires a number: the gap between one word and the next. In March twenty twenty-four MLCommons — the multi-vendor consortium behind MLPerf — set that gap at two hundred milliseconds for a chat benchmark, and published its reasoning: roughly two hundred and forty words a minute, often cited as average reading speed. A year later, for an interactive variant, forty. This March, for a reasoning workload, fifteen. Three different benchmarks, and after the first the number stopped being aimed at a reader.

There are two live answers to the half that is forced.

The first accepts the sequence and attacks the waiting: **speculative decoding**. A small cheap model guesses the next several pieces; the big model checks them all in one pass — which it can, because checking is the parallel half. Guesses that survive the check cost no further big-model pass; the rest are thrown away.

What makes it more than a shortcut is that the output distribution is provably unchanged. Yaniv Leviathan, Matan Kalman and Yossi Matias at Google Research posted it in November twenty twenty-two, reporting two to three times faster; a DeepMind team published the same guarantee two months later, at two to two and a half times, theirs stated as exact within hardware numerics.

Both are authors measuring their own method, the first at batch size one — one user, one machine, nothing else running. Last December a Berkeley group with no stake in it measured the technique on a production stack, and said plainly those figures come from an unrealistic setting. What they got: one point nine six times for a single user, falling to one point two one when a hundred and twenty-eight requests share the machine.

The second answer looks like it refuses the sequence: generate a block at once and refine it, the way image models do. But when Google released a diffusion text model in June, its own model card called it *block-autoregressive* — past two hundred and fifty-six pieces it finishes a block, commits it, and starts the next conditioned on everything committed. The loop did not go away; the stride got wider. And Google's recommendation, about its own release:

> For applications that demand maximum quality, we recommend deploying standard Gemma 4.

> — *Google, 'DiffusionGemma: 4x faster text generation', The Keyword, 10 June 2026, section 'Unlocking new value for developers', under 'Experimental status & production recommendations'; https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/, read 6 September 2026*

Two things, each on a page, on a date.

One. That fifteen-millisecond MLPerf scenario is the first benchmark rule anywhere to require speculative decoding. In its first round the results file carries three entries, from two organisations, all on one vendor's chips. Watch the next file for a third organisation, or a result on somebody else's hardware.

Two. Three days ago OpenAI shipped mid-turn steering: you can talk to a response while it is still being written. Its documentation states the limit in the same breath.

> Steering does not rewrite output already sent to your application, undo earlier actions, or cancel tools that have already started.

> — *OpenAI, 'Mid-turn steering', OpenAI API documentation, opening paragraphs; https://developers.openai.com/api/docs/guides/steering, read 6 September 2026*

Today's subject, written as a product limitation by the company shipping it. If anyone can ever take back a word already sent, that sentence changes first.

One idea, in two halves.

**It writes one piece at a time. It reads everything, every time.** The first half explains the moving cursor, the wait before the first word, and a whole industry of tricks for getting round it. The second half is what the word autocomplete hides — and there is a number on it, because a two-token window turns the same machine into a predictor that disagrees with itself four times in five.

An old shape: a man counting vowels in Pushkin by hand in nineteen thirteen, another opening books at random in nineteen forty-eight and stopping because the labour became enormous. What fills it in, and what it costs, have changed. The shape has barely moved.

Both are free to read: Shannon's paper, where section three is the ladder, and Jurafsky and Martin's *Speech and Language Processing*.

Tomorrow: when the model reads everything in front of it, how does it work out which parts matter? That mechanism has a name you have heard, and it is what made the current generation of these systems possible.

That was day three. Thank you for listening.

---

## Sources (16)

- A. A. Markov — *An Example of Statistical Investigation of the Text Eugene Onegin* — lecture 23 Jan 1913; translation Dec 2006
- Brian Hayes — *First Links in the Markov Chain* — Mar–Apr 2013
- Claude Shannon — *A Mathematical Theory of Communication* — Jul & Oct 1948
- Jurafsky & Martin — *Speech and Language Processing*, 3rd ed. draft — released 19 Aug 2026
- Direct measurement, one desktop machine — 6 Sep 2026
- Zhicheng Lin — *Six misconceptions about large language models* — 19 Aug 2026
- Gary Marcus — *The Dark Risk of Large Language Models* — 29 Dec 2022
- Anthropic (Lindsey et al.) — *On the Biology of a Large Language Model* — 27 Mar 2025
- OpenAI — *Streaming API responses*; *Mid-turn steering* — read 6 Sep 2026; steering shipped 3 Sep 2026
- OpenAI — `openai-python` initial commit `3c6d4cd6`, Greg Brockman — 25 Oct 2020
- MLCommons — Llama 2 70B benchmark; Inference v5.0; GPT-OSS/DeepSeek-R1 update; v6.0 results file — 27 Mar 2024; 2 Apr 2025; 24 Mar 2026; 1 Apr 2026
- Leviathan, Kalman & Matias — speculative decoding — 30 Nov 2022
- Chen et al. — speculative sampling — 2 Feb 2023
- Liu, Yu, Park, Stoica & Cheung — *Speculative Decoding: Performance or Illusion?* — 31 Dec 2025
- Google — *DiffusionGemma: 4x faster text generation*; `google/diffusiongemma-26B-A4B-it` model card — 10 Jun 2026; card read 6 Sep 2026
- Artificial Analysis — DiffusionGemma providers — read 6 Sep 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
