# Understanding AI in a Month 12: Why Models Think Longer

Understanding AI in a Month — Day 12 · 2026-09-21

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

Open Anthropic's developer documentation today and there's a table, model by model, of whether thinking can be switched off. Five of the current ones are marked always on. Underneath it, one sentence.

> Models marked Always on cannot turn thinking off.

> — *Anthropic, 'Troubleshooting thinking', Claude Platform documentation, the per-model thinking support table; read 21 September 2026*

Ask one of those models not to think and the request comes back an error. At OpenAI, the changelog entry for the third of September says its newest model doesn't accept the setting called none — the one meaning don't think at all. Google's documentation, updated on the seventeenth of September, lists twelve models, eleven of them thinking by default.

Two years ago, a model that worked through a problem before answering was a separate product you chose on purpose. Now it's the floor, and on the best models you can't get off it.

Yesterday was what a benchmark score proves. Today is the thing that happens between your question and the answer.

Two readings of that, and both go too far.

The first is that the panel of thinking you can watch is a window into the machine's mind. What you can watch is not that, and all three companies say so in their own documentation. OpenAI decided in twenty twenty-four not to show raw chains of thought, only a summary. Anthropic's documentation says what you see is never the raw chain of thought, and Google says the same — adding that you are charged for the full thoughts, not the summary you see. Whether what the machine does deserves the word thinking is a different argument, and none of this settles it.

The second reading is the cynical one: it's padding, you're billed by the word, and the word count is the product. That one is answerable with measurements — and the measurements are stranger than either side expects.

The idea is called test-time compute: work done when the question is asked, not when the model was trained. The cleanest way in is a paper from nineteen fifty.

In March nineteen fifty, Claude Shannon published the first design for playing chess on a general-purpose computer. He described a strategy that examines every legal continuation to a fixed depth and then judges the positions — he called it type A — and then said what was wrong with it.

> A good human player examines only a few selected variations and carries these out to a reasonable stopping point.

> — *Claude E. Shannon, 'Programming a Computer for Playing Chess', Philosophical Magazine ser. 7, 41(314), March 1950, section 6 'Improvements in the Strategy'*

So he proposed a second strategy, type B: pick which lines are worth following, and follow those as far as they stay useful. Same knowledge of chess; different way of spending the time you have when the move is played. That is the question every effort setting is a version of.

Now the part that makes it apply to a language model. In November twenty twenty-one Maxwell Nye and colleagues wrote the constraint down exactly.

> Given a fixed number of layers and a fixed amount of computation time, the model cannot adapt the amount of compute spent on a problem to its difficulty before producing an output.

> — *Maxwell Nye and colleagues, 'Show Your Work: Scratchpads for Intermediate Computation with Language Models', arXiv 2112.00114 v1, 30 November 2021, introduction*

A model does the same amount of arithmetic for every token it produces. It cannot sit and think quietly, because there's nowhere quiet to sit. So if it's going to do more work on your question, that work has to come out as text. Their fix was to let it write the intermediate steps into what they called a scratchpad first, and models that had failed at long addition started getting it right. Two months later Jason Wei and colleagues at Google got the same effect from the prompt alone — though their first version warned that below roughly a hundred billion parameters they got fluent but illogical chains of thought.

So: thinking tokens are ordinary tokens. The model buys computation by writing. Everything else follows from that, and two things in particular.

The first is that writing more gives you several attempts, and now you have to choose one. In October twenty twenty-one Karl Cobbe and colleagues at OpenAI published the answer.

> At test time, we generate many candidate solutions and select the one ranked highest by the verifier.

> — *Karl Cobbe and colleagues (OpenAI), 'Training Verifiers to Solve Math Word Problems', arXiv 2110.14168, abstract, 27 October 2021*

A verifier is a second model trained to score a solution: generate a hundred answers, rank them, keep the best.

The second is that a verifier grades the answer — and you could grade the working instead. In twenty twenty-three a team at OpenAI led by Hunter Lightman put the choice like this.

> To train more reliable models, we can turn either to outcome supervision, which provides feedback for a final result, or process supervision, which provides feedback for each intermediate reasoning step.

> — *Hunter Lightman and colleagues (OpenAI), 'Let's Verify Step by Step', arXiv 2305.20050 v1, abstract, 31 May 2023*

They paid for eight hundred thousand human judgements of single steps. Then, on five hundred held-out competition maths problems, choosing among one thousand eight hundred and sixty candidate solutions each, the step-trained ranker picked a correct answer seventy-eight per cent of the time, against seventy-two for the answer-trained ranker and seventy for a plain majority vote. Every one of those conditions matters, which is why I said them all.

That's the toolkit. The history is how it moved from outside the model to inside.

Nineteen fifty, Shannon. Twenty sixteen, AlphaGo, whose paper reports that without any lookahead search its networks alone played Go about as well as the best searching programs — and with search on top it beat the European champion five-nil. Instinct and deliberation came apart.

Then the language-model versions: scratchpads and verifiers in twenty twenty-one, prompted steps and majority voting in twenty twenty-two, graded working in twenty twenty-three. All of it done to the model from outside — a prompt, a sampler, a ranking step.

September twenty twenty-four is when it moved inside. OpenAI released a model trained to produce the working-out itself, and described how.

> Our large-scale reinforcement learning algorithm teaches the model how to think productively using its chain of thought in a highly data-efficient training process.

> — *OpenAI, 'Learning to Reason with LLMs', 12 September 2024, the opening section, before 'Evals'; Internet Archive capture of 13 September 2024 of https://openai.com/index/learning-to-reason-with-llms/*

The company said it had found performance improving both with more reinforcement learning in training and with more thinking at the question — and that the limits on this differed from pre-training's, and were still under investigation. The concrete version, from the same post: on that year's American maths olympiad qualifier, the older model solved about twelve per cent of the problems. The new one averaged seventy-four per cent with one attempt, eighty-three taking the majority answer of sixty-four attempts, and ninety-three when a thousand attempts were re-ranked by a learned scorer.

Four months later DeepSeek published a chart nobody had shown before: over the course of reinforcement learning, its model's answers got longer on their own. Nobody set a length, and nothing rewarded length: the rewards were for right answers and for putting the working in the right place. DeepSeek calls the growth intrinsic to the model.

Two things happened next, and they point in opposite directions.

First, the measurements, and they break the obvious story. On the third of September the ARC Prize Foundation — which owns the benchmark, not the model — ran OpenAI's newest at six effort settings and reported this.

> Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing the total number of model calls and tokens.

> — *ARC Prize Foundation (Greg Kamradt), 'OpenAI's GPT-6 Astra on ARC-AGI-3', 3 September 2026, section 'Astra Results'*

At the highest setting it scored sixty-two point seven per cent and the run cost twenty-six thousand dollars. At medium it scored thirty-eight point six and cost forty-eight thousand. More thinking per move, fewer moves, less money — generally, which is the foundation's own word, because neither column runs in order. Low cost less than medium; low also scored seventeen point five, below the setting called none at thirty-five point two. There are no error bars on that table, so it shows the rungs don't climb in a straight line — not that one rung is worse than another.

The evaluator Artificial Analysis measured the same model at five effort settings on its own index. Lowest to highest, it spent about eighteen times the reasoning tokens to gain seven points, at four times the money — and the wait before the first word went from three seconds to over four minutes.

Second, the two companies moved opposite ways inside a fortnight. On the first of September Anthropic shipped models whose thinking cannot be switched off. On the fourteenth, OpenAI's release notes said this about its consumer product.

> We’re retiring automatic switching from Instant to Thinking (reasoning) for ChatGPT Plus and Pro users globally.

> — *OpenAI, 'ChatGPT Release Notes', the entry 'Changes to automatic switching to thinking in ChatGPT (Plus and Pro)' dated 14 September 2026; Internet Archive capture 20260918133646*

On the developer side the off switch is disappearing; on the consumer side the choice is going back to the person typing. Both dated, both in the companies' own words.

Three things you can check on a page.

One. Anthropic's per-model thinking table. Today five models are marked always on, two more on with an opt-out. When the next model appears, read which column its row lands in.

Two. OpenAI's changelog at the next model release. Its entry for the third of September says in so many words that the newest model does not support the none setting. One line tells you whether that option comes back or spreads.

Three, and it's the interesting one. OpenAI's guide says that model rejects the none setting. ARC Prize's table of the third of September has a none row for it anyway, with a score and a cost. Both are published, neither explains the other. Watch whichever moves first.

The idea to keep is test-time compute: the work a system does on your question after its training is finished.

Hold it together with the mechanism, because the mechanism keeps you honest. There is no separate thinking organ: every token of working-out is a token the machine generated, and you pay for them. OpenAI's guide puts it plainly.

> While reasoning tokens are not visible via the A P I, they still occupy space in the model’s context window and are billed as output tokens

> — *OpenAI, 'Reasoning models', OpenAI API documentation, section 'How reasoning works'; https://developers.openai.com/api/docs/guides/reasoning, read 21 September 2026*

> As printed in the source: “While reasoning tokens are not visible via the API, they still occupy space in the model’s context window and are billed as output tokens”

Anthropic says the same, and Google's price list has a row labelled output price including thinking tokens.

So any claim about a reasoning model raises a pair of questions. Does more thinking buy anything on this task — and what did it cost? Separate answers, and this year they stopped moving together.

To read more, the encyclopedia has articles on test-time compute, reasoning models, and process supervision.

Tomorrow: how a smaller or sparser system inherits capability, and what a unit of quality costs.

That was day twelve. Thank you for listening.

---

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
