# Understanding AI in a Month 16: What One Answer Costs

Understanding AI in a Month — Day 16 · 2026-09-26

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

On Tuesday, the research group Epoch AI published a report called The plunging price of thought. Its finding: the cost of reaching a given level of performance on its benchmarks has fallen about forty-seven per cent a quarter since twenty twenty-three, which compounds to about thirteen times a year. One example it gives: OpenAI's o3, released in January last year, could reach seventy-five per cent on a multiple-choice exam of PhD-level science at an estimated thirty cents a question. Just under eighteen months later, GPT-5.6 Luna matched that score for four hundredths of a cent. That's one pair, not the average. And Epoch is careful about who actually collects the savings. It tracks the cheapest model that can reach each score, which means, in its words,

> we implicitly posit an AI user who relentlessly searches for the most cost-effective model for each task, when real users do not switch models so often, and therefore do not reap quite the same savings.

> — *Epoch AI, Luke Emberson and David Roodman, 'The plunging price of thought', report published 22 September 2026, section 'Overview', paragraph on the report's limitations; https://epoch.ai/publications/the-plunging-price-of-thought, read 26 September 2026*

Today: what it takes to produce one answer, and why its price comes in half a dozen rates.

Two readings of that headline mislead.

The first is that the AI you use got thirteen times cheaper this year. Epoch's number is for a fixed level of performance, bought from whichever model is now cheapest. A team at MIT, Hans Gundlach and colleagues, in a paper revised in March, also found that kind of price falling, five to ten times a year. But they estimated that the price of running the frontier models themselves is rising, three to eighteen times a year, because the models are bigger and do more reasoning. Both can be true. And a price can rise with no change to the model: Google launched Gemini 3.8 Flash on the second of this month at an introductory price that doubles on the first of January.

The second reading is that a token has a price. OpenAI's middle model, GPT-6 Sol, has at least six on its price sheet today. Two dollars per million tokens you send. Ten dollars per million it writes. Twenty cents per million for input it has already seen. Double the input rate once a prompt passes two hundred and seventy-two thousand tokens. Half price if you'll wait. Double if you want it faster. Yet in arithmetic, reading a token and writing one cost about the same: Google engineers counted two operations per parameter for every token, either way. Writing is billed at five times the rate. Why is today's idea.

Every answer has two phases, and they behave nothing alike.

The first is reading your prompt. The whole prompt goes through the model at once, every token side by side, and the chip's arithmetic is kept busy. That phase is called prefill, and it largely decides how long you wait for the first word.

The second is writing, called decode. The answer comes out one token at a time, and for each one the chip reads the model's weights out of memory again. Two episodes ago we saw that for a single user it's that reading, not the arithmetic, that sets the pace.

And decode has a second thing to read. To choose each new token, the model looks back at every token before it. Rather than redo that work every time, it keeps, for every earlier token and at every layer, two sets of numbers called keys and values. That store is the K V cache. In twenty nineteen Noam Shazeer at Google put the problem it creates like this.

> the speed of incremental Transformer inference on modern computing hardware is limited by the memory bandwidth necessary to reload the large "keys" and "values" tensors which encode the state of the attention layers.

> — *Noam Shazeer (Google), 'Fast Transformer Decoding: One Write-Head is All You Need', arXiv 1911.02150 v1, 6 November 2019, section 1 'Introduction', second paragraph*

How big does it get? On this show's arithmetic, from the published configuration of Alibaba's open Qwen3 model with thirty-two billion parameters, every token of a conversation takes about a quarter of a megabyte of cache. A conversation of thirty-two thousand tokens, the model's full native length, needs about eight and a half gigabytes, for one user.

Now batching. Since every decode step reads the whole model anyway, a server makes each pass produce the next token for many conversations at once. The weights are read once and shared. The cache isn't. A Google team spelled that out in November twenty twenty-two, in a paper called Efficiently Scaling Transformer Inference. Its authors are the ones who called the two phases prefill and decode.

> (unlike the weights) the K V cache is unique for each sequence in the batch.

> — *Reiner Pope, Sholto Douglas, Aakanksha Chowdhery and colleagues (Google), 'Efficiently Scaling Transformer Inference', arXiv 2211.05102 v1, 9 November 2022, section 2 'Inference Cost Tradeoffs', end of the 'Compute costs' paragraph*

> As printed in the source: “(unlike the weights) the KV cache is unique for each sequence in the batch.”

That year a system called Orca, from Seoul National University and the company FriendliAI, re-formed the batch at every step, so a finished answer leaves and a new one joins without waiting for the rest. The next year a team led from Berkeley built vLLM, which stores the cache in small pages, the way an operating system handles memory. They measured the systems before it using only about a fifth to two-fifths of their cache memory for actual tokens.

And quantisation: storing each weight in fewer bits. Half the bytes means half as much to read at every step, and room for more conversations. In twenty twenty-two Tim Dettmers and colleagues reported converting a model of a hundred and seventy-five billion parameters to eight bits without losing performance, using a method that keeps a small set of outlier values at sixteen.

Put the four together on one of NVIDIA's eighty-gigabyte H100 chips, with Qwen3's weights at eight bits. Again on this show's arithmetic, counting only memory: that leaves room for about forty conversations of four thousand tokens, or five of thirty-two thousand. One user alone could get about a hundred tokens a second. Forty users sharing each pass get about forty each, but well over fifteen hundred between them. Five long conversations get about two hundred between them.

That's the pull. Batch more, and total output rises while each user waits longer for every word. Let conversations grow, and fewer fit, so the total falls. Make one user fast, and each token takes more of the chip. The research firm SemiAnalysis, describing its serving benchmark, puts the first part plainly.

> Large batches enable better G P U utilization and higher token throughput, but they split available resources across more requests, slowing down token processing per user.

> — *SemiAnalysis, 'InferenceMAX' launch post, SemiAnalysis newsletter, 9 October 2025, section 'The Fundamental Trade-off between Throughput & Latency/Interactivity', second paragraph; https://newsletter.semianalysis.com/p/inferencemax-open-source-inference, read 26 September 2026*

> As printed in the source: “Large batches enable better GPU utilization and higher token throughput, but they split available resources across more requests, slowing down token processing per user.”

Epoch measures the fall. An earlier Epoch analysis named two well-known reasons: models getting smaller, and hardware more cost-effective. The machinery under one answer changed too, and it can be dated. Shazeer named the memory problem in twenty nineteen and proposed a smaller cache, with a layer's attention heads sharing one set of keys and values. Continuous batching, eight-bit weights for the largest models and the Google paper came in twenty twenty-two; paged cache memory in twenty twenty-three. From late twenty twenty-three, teams at the University of Washington and Microsoft, then Peking University, proposed putting prefill and decode on different machines, since one phase wants arithmetic and the other memory; Moonshot AI has described serving its Kimi assistant that way. In May twenty twenty-four DeepSeek reported an attention design that, against an earlier model of its own, cut the cache by ninety-three per cent.

Last week the benchmark consortium MLCommons published its latest inference results. Its working-group chairs found the typical result per chip, on one test run for six rounds, up about five and a half times. They gave three reasons: some submissions now use four-bit numbers where earlier rounds used eight, within the benchmark's accuracy rules; newer chips; and better software, even on the same chips. The results are vendors' own submissions, under rules the other submitters review. Two of those three reasons are this episode's subject: fewer bits per number, and serving software.

Now read a price sheet with the two phases in mind. None of the sheets says why output costs more, so what follows is this show's reading. Output at five times input is decode: one token per pass. Cached input at a tenth is prefill that doesn't have to be redone. OpenAI's guide says caching lets it

> Avoid recalculating a prompt prefix that the model has already processed.

> — *OpenAI, 'Prompt caching' guide, OpenAI API documentation, section 'Why prompt caching matters', first listed benefit ('Compute-efficient'); https://developers.openai.com/api/docs/guides/prompt-caching, read 26 September 2026*

A kept cache has to be stored, and Google charges for that: four dollars fifty per million tokens per hour on Gemini 3.1 Pro. Half price for batch work is a customer letting the provider fill its batches. And in July OpenAI renamed its priority tier Fast mode: up to two and a half times faster, at twice the price.

Where providers differ most is long context. OpenAI charges more past two hundred and seventy-two thousand tokens, Google past two hundred thousand, and xAI, from two hundred thousand, charges every token of the request at the higher rate. DeepSeek offers a million-token window with no long-context tier at all, and halves its prices outside its peak hours. DeepSeek doesn't say why, and a price isn't a cost. But its published designs since twenty twenty-four have been built to shrink the cache.

One. The twenty-first of November. OpenAI's sheet says a promotional price for its older GPT-5.6 Sol lasts at least until then. Watch whether that row changes.

Two. The first of January. Gemini 3.8 Flash's introductory price ends: input from seventy-five cents per million tokens to a dollar fifty, output from three seventy-five to seven fifty. Same model, twice the price.

Three. MLCommons says a new benchmark, MLPerf Endpoints, will replace its data-centre inference benchmark, and gives no date. When results appear, watch whether they report speed per user beside output per chip.

The idea to keep is the K V cache: the weights are shared by everyone in the batch, and the cache belongs to one conversation. On this show's reading, most of the rates on a price sheet line up with one side of that or the other.

So when you see the price of an answer, ask three things. Is this token being read, re-read from a cache, or written? How long is the conversation behind it? And how fast does it have to arrive?

To read more: Reiner Pope and colleagues at Google, Efficiently Scaling Transformer Inference, twenty twenty-two.

---

## Sources (26)

- Epoch AI (Luke Emberson, David Roodman), *The plunging price of thought* — 22 Sep 2026
- Epoch AI (Ben Cottier and colleagues), *LLM inference prices have fallen rapidly but unequally across tasks* — 12 Mar 2025
- Gundlach, Lynch, Mertens, Thompson, *The Price of Progress*, arXiv 2511.23455 v2 — 23 Mar 2026
- OpenAI, API pricing page; prompt caching, Fast mode and Flex guides — read 26 Sep 2026
- Google, Gemini Developer API pricing; Gemini 3.8 Flash announcement — 24 Sep 2026; 2 Sep 2026
- xAI, API pricing page — 21 Sep 2026
- DeepSeek, Models & Pricing — read 26 Sep 2026
- Z.ai, pricing page — read 26 Sep 2026
- Pope et al. (Google), *Efficiently Scaling Transformer Inference*, arXiv 2211.05102 — 9 Nov 2022
- Shazeer (Google), *Fast Transformer Decoding*, arXiv 1911.02150 — 6 Nov 2019
- Ainslie et al. (Google), *GQA*, arXiv 2305.13245 — May 2023
- DeepSeek-AI, *DeepSeek-V2*, arXiv 2405.04434 — May 2024
- Yu et al., *Orca*, OSDI 2022 — Jul 2022
- Kwon et al., *PagedAttention / vLLM*, arXiv 2309.06180 (SOSP 2023) — 12 Sep 2023
- Vanhoucke, Senior, Mao (Google), *Improving the speed of neural networks on CPUs* — 2011
- Jacob et al. (Google), arXiv 1712.05877 — 15 Dec 2017
- Dettmers et al., *LLM.int8()*, arXiv 2208.07339 — 15 Aug 2022
- Frantar et al., *GPTQ*, arXiv 2210.17323 — 31 Oct 2022
- Patel et al., *Splitwise*, arXiv 2311.18677 — 30 Nov 2023
- Zhong et al., *DistServe*, arXiv 2401.09670 — 18 Jan 2024
- Agrawal et al., *Sarathi-Serve*, arXiv 2403.02310 — 4 Mar 2024
- Qin et al. (Moonshot AI, Tsinghua), *Mooncake*, arXiv 2407.00079 — 24 Jun 2024
- SemiAnalysis, InferenceMAX launch post — 9 Oct 2025
- MLCommons, MLPerf Inference v6.1 results and chairs' analysis — 16–17 Sep 2026
- Alibaba Qwen team, Qwen3-32B model card and config; DeepSeek-V3 and Qwen3-235B-A22B configs — read 26 Sep 2026
- NVIDIA, H100 product page — read 26 Sep 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
