# Understanding AI in a Month 5: How an Answer Unfolds

Understanding AI in a Month — Day 5 · 2026-09-09

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

A year ago this week, a lab called Thinking Machines published a small, patient experiment. One open-weight model built by somebody else, one prompt — *tell me about Richard Feynman* — and the randomness turned all the way down: setting zero, which is supposed to mean the same answer every time.

Then they asked a thousand times.

They got eighty different answers, and the commonest turned up seventy-eight times — so more than nine runs in ten produced something other than the usual reply. All thousand agreed for the first hundred and two pieces of text — then nine hundred and ninety-two said Feynman was born in "Queens, New York", and eight said "New York City".

The first thing people say is that there is a dial marked randomness, and that zero makes the thing repeatable. Here is Anthropic's API reference, writing to the people building on it.

> Note that even with temperature of zero point zero the results will not be fully deterministic.

> — *Anthropic, Messages API reference, 'Create a Message', body parameter 'temperature'; https://platform.claude.com/docs/en/api/messages/create, read 8 September 2026*

> As printed in the source: “Note that even with `temperature` of `0.0`, the results will not be fully deterministic.”

The second is four words long: *it forgot my conversation*. Those words name no mechanism. The text may have been dropped before it reached the model, replaced by a shorter stand-in, or sent in full and used badly — and nothing inside the model decides any of that. I will show you who does.

And a debt from yesterday, which ended with the model scoring every piece of text it could say next — tens of thousands at once — and stopped there on purpose. A score is not a word. Something has to pick one.

Two ideas: how you pick, and how much you may look at while picking.

**First, the picking.** The simplest rule takes the highest-scoring next token every time. That is called greedy, and the textbook behind this series — Jurafsky and Martin, in the draft they released in August — is exact about what it buys you.

> Indeed, greedy decoding is so predictable that it is deterministic; if the context is identical, and the probabilistic model is the same, greedy decoding will always result in generating exactly the same string.

> — *Daniel Jurafsky and James H. Martin, 'Speech and Language Processing', 3rd edition draft of 19 August 2026, chapter 7, section 7.6.1 'Greedy decoding', p. 21; https://web.stanford.edu/~jurafsky/slp3/7.pdf*

Notice the two conditions. The Feynman run held both fixed and still came back eighty ways — so the second one is about the arithmetic, not the model.

The same page says greedy decoding is not, in practice, used with large language models — and why.

> In other words, greedy decoding is too boring, and random sampling is too random.

> — *Daniel Jurafsky and James H. Martin, 'Speech and Language Processing', 3rd edition draft of 19 August 2026, chapter 7, section 7.6.2 'Random sampling', p. 22; https://web.stanford.edu/~jurafsky/slp3/7.pdf*

Too boring, because always taking the likeliest word gives text the textbook calls generic and often quite repetitive. Too random, because the obvious alternative — turn the scores into probabilities and draw one — reaches into a tail of unlikely words that get picked often enough to wreck a sentence.

Two ways to change it. You can **reshape** the odds first: sharpen them and you approach greedy, flatten them and stranger things get through. That is the dial you have heard of, and the textbook traces the idea to thermodynamics.

Or you can **cut** them. In twenty eighteen Angela Fan and colleagues at Facebook, generating fiction from writing prompts, drew from only the ten likeliest candidates. Completely random sampling, they wrote, can introduce very unlikely words, and the model has not seen such mistakes at training time. Holtzman and colleagues, two years later, named what is wrong with a fixed count: set it small and you risk bland text, set it large and you let in candidates that do not belong. Take a fixed *share* instead, and the count moves by itself.

I measured that here, on five hundred words of prose written for the purpose and never published, counting at every position how many candidates cover ninety per cent of the probability. The middle was **fourteen**; the smallest, one; the largest five thousand five hundred and seventy. In one passage.

**The second idea is the bound.** There is a ceiling on how much text can be in front of the model at once, and it is called the context window. The textbook credits sampling from a language model to Shannon in nineteen forty-eight and to Miller and Selfridge in nineteen fifty. Day three did Shannon. The second name is today's: Miller and Selfridge wrote no computing paper at all. Theirs was *Verbal context and the recall of meaningful material*, filed by the National Library of Medicine under memory and recall — the oldest citation under this machinery is an experiment on people.

And one word to carry. When you cut the odds before drawing, the textbook's verb is **truncate**. It will come back, meaning something else entirely.

So where did the famous million come from? Not this year.

The fifteenth of February, twenty twenty-four. Google announced Gemini one point five Pro and said it had run up to a million pieces of text consistently — the longest window, it said, of any large-scale foundation model yet. Google's claim about Google, read in a copy of the page taken the afternoon it went up.

Every qualifier is Google's own: a hundred and twenty-eight thousand was the standard window that day, and the million went to a limited group in private preview, labelled experimental, with longer waits.

Fourteen months later OpenAI brought the same number to its own developers, saying it had trained the model to reliably attend across the full million. Then, same page, same day, it published scores on harder tests it had built itself — because, in its own words, few real-world tasks are as straightforward as retrieving a single obvious planted answer. On a graph-search puzzle: below a hundred and twenty-eight thousand pieces of text, about sixty-two per cent; above it, nineteen. A bucket rather than a measurement at a million — but it is the company selling the million publishing the number that complicates it.

Anthropic followed: a million in beta for one model in August twenty twenty-five, out of beta on newer ones this March. And a week ago Google announced a new model, and the announcement does not mention the context window once. In twenty twenty-four it was the headline.

Two things, and the first surprised me.

**The window went backwards.** On the thirtieth of April this year Anthropic retired the million-token option on two of its older models; the release note says requests over the standard limit now return an error. A model that could take a million pieces of text on the twenty-ninth of April could not on the first of May. The release note records a preview ending, not a discovery.

**Second: the model is not the one deciding.** When the text you send is already too long the input is refused, and Anthropic's documentation adds a phrase worth hearing: *on every model*. That is a rule written into an API, not a model running out of room.

Everything else is software above the model deciding, and the decisions differ inside one company. In OpenAI's own specification its Responses API can drop items from the beginning of a conversation, and its Assistants API, retired in August, dropped messages from the middle instead. A third replaces the old material with a shorter stand-in — and here is OpenAI's own guide on that.

> It is opaque and not intended to be human-interpretable.

> — *OpenAI, 'Compaction', OpenAI API documentation, section 'Server-side compaction'; https://developers.openai.com/api/docs/guides/compaction, read 8 September 2026*

"It forgot" names no actor. If the text was dropped, something above the model dropped it and somebody chose which end — and in one case what sits in its place is sealed.

Now the contrast. That experiment came with a fix, and it works where it was tested: vLLM and SGLang have both shipped it, off unless you turn it on. In August the authors of a competing design, CoRun, priced it at more than double the latency and up to three-quarters of the throughput. And on the third of September Zhu and Zhang's audit found identical inputs still ranked differently across four providers — and found the fix, switched on, held only while the server was quiet.

Diagnosed, remedied, remedy shipped switched off, and still not simply cured.

Two things, each on a page, on a date.

One. The vLLM project documents this on a page whose footer still reads the seventh of September.

> Batch invariance is currently in beta.

> — *vLLM, 'Batch Invariance', vLLM documentation (page dated 7 September 2026), the note at the head of the page; https://docs.vllm.ai/en/latest/features/batch_invariance.html, read 8 September 2026*

Its own note says switching it on may cost performance against the default, deliberately. Watch two words: *beta*, and *default*. Beta is whether the thing is finished; default is whether anyone has to ask for it. They move separately, and either would be news.

Two. A long-context leaderboard called Context Arena, run by one person, Dillon Uzar. It knows of five hundred and forty-five models; seventy-three appear on its eight-needle board; and at the longest lengths — half a million pieces of text and up — **nine** have any score at all, seven of them from one company. I counted that this morning from the site's own data. The best of the nine scores a flat hundred per cent at four to eight thousand, and sixty-three and a half in its longest band.

**The number is a capacity, not a competence.** Both of today's mechanisms are choices, not properties. Something chooses how the next piece is drawn; something chooses what stays in front of the model. Neither is the machine deciding, and both are written down somewhere, dated.

And some of those choices have stopped being yours. On Anthropic's Sonnet 5 since June, and on OpenAI's GPT-6 Astra since the third of September, changing that first dial away from its default returns an error.

If you keep one image, keep the desktop in this room. One prompt, greedy, randomness off. At a batch of one: the same answer, five runs out of five. At a batch of sixty-four: the same answer five out of five — a *different* answer.

Nothing in there was thinking differently. Only the batch size was deliberately changed — and by the lab's account, that changes the order some numbers get added in. On your own computer you control that; on somebody's server you do not. The lab that found it put it best.

> From the perspective of an individual user, the other concurrent users are not an input to the system but rather a nondeterministic property of the system.

> — *Horace He and Thinking Machines Lab, 'Defeating Nondeterminism in LLM Inference', 10 September 2025, doi:10.64434/tml.20250910, section 'Batch invariance and “determinism”'; https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/*

> As printed in the source: “From the perspective of an individual user, the other concurrent users are not an “input” to the system but rather a nondeterministic property of the system.”

Tomorrow: we give the model search, files, tools and a memory, and find out how much of what you thought was the model never was.

That was day five. Thank you for listening.

---

## Sources (22)

- Horace He and Thinking Machines Lab — *Defeating Nondeterminism in LLM Inference* — 10 Sep 2025
- Jurafsky and Martin — *Speech and Language Processing*, 3rd edition draft — released 19 Aug 2026
- Claude Shannon — *A Mathematical Theory of Communication* — Jul and Oct 1948
- George A. Miller and Jennifer A. Selfridge — *Verbal context and the recall of meaningful material* — Apr 1950
- Angela Fan, Mike Lewis, Yann Dauphin — *Hierarchical Neural Story Generation* — 13 May 2018 (v1)
- Holtzman, Buys, Du, Forbes, Choi — *The Curious Case of Neural Text Degeneration* — ICLR 2020 (preprint 22 Apr 2019)
- Google — *Our next-generation model: Gemini 1.5* — 15 Feb 2024
- OpenAI — *Introducing GPT-4.1 in the API* — 14 Apr 2025
- Anthropic — Claude Platform release notes — entries 12 Aug 2025 – 3 Sep 2026
- Anthropic — Messages API reference, Context windows, Compaction — read 8 Sep 2026
- OpenAI — `openai/openai-openapi` specification and the compaction guide — read 8 Sep 2026
- Google — Gemini API long-context guide — updated 22 Jun 2026
- Google — Gemini 3.8 Flash announcement — 2 Sep 2026
- vLLM — Batch Invariance documentation — page dated 7 Sep 2026
- SGLang — Deterministic Inference documentation — read 8 Sep 2026
- Zhao and colleagues — *CoRun: Padding is Simple and Efficient for Deterministic LLM Inference* — 14 Aug 2026 (v1)
- Joshi, Aggarwal, Das and colleagues — 12 Jan 2026 (v1)
- Zhu and Zhang — preregistered reliability audit — 3 Sep 2026 (v1)
- Aditi Patodiya — *Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving* — 4 Sep 2026 (v1)
- HELMET — Princeton PLI and Intel — ICLR 2025
- Context Arena, run by Dillon Uzar — read 9 Sep 2026
- Direct measurement, one desktop machine — 8 Sep 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
