# Understanding AI in a Month 6: What the Application Adds

Understanding AI in a Month — Day 6 · 2026-09-10

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

In May twenty twenty, a group at Facebook AI Research and two universities published a paper on a system that could look things up before answering. Near the end there is a small experiment.

They took eighty-two heads of state who had changed between twenty sixteen and twenty eighteen, and asked about each with the same question, the office filled in. *Who is the prime minister of the United Kingdom?*

Then they ran it twice. Nothing retrained, same system, same questions. The only difference was which copy of Wikipedia it could consult: one from December twenty sixteen, one from twenty eighteen.

Start with the failure, because it teaches more. Handed the copy from the wrong period, it got four answers in a hundred right; the other way round, twelve. Handed the copy that matched, about seven in ten.

Same machine, different filing cabinet — and only twenty-one per cent of what it said came out the same either way.

Two readings, and I want to head off both.

The first is the cheerful one: looking things up makes you right. The four per cent answers that. Whatever it did with the other ninety-six, it was not declining to answer. A wrong drawer is worse than no drawer.

And not a museum piece. In twenty twenty-four a Stanford-led team ran what it calls the first pre-registered evaluation of commercial legal research tools — three products, from two companies, in a market where vendors had advertised citations free of invention. Journal of Empirical Legal Studies, April twenty twenty-five.

> We demonstrate that the providers' claims are overstated.

> — *arXiv 2405.20362 v1 (Magesh et al.), abstract*

> As printed in the source: “We demonstrate that the providers’ claims are overstated.”

They measured those products making things up between seventeen and thirty-three per cent of the time — and the same sentence carries the other half: still better than a general chatbot. Better. Not what the page said.

The second reading is the one today is most likely to cause: that the model barely matters and the software around it is what does. Hold that — the best evidence I found refuses it, and I would rather you heard the refusal from me.

Yesterday ended with something above the model deciding what stays in front of it. Today the same layer decides what goes in — and, the half yesterday never reached, it is also the only thing that can act.

One boundary, and everything hangs off it.

**A model reads text and writes text. That is the whole of what it does.** Everything else a product appears to do — search, open a file, run something, remember you — is ordinary software choosing what text to put in front of it, and deciding what to do with what it hands back.

Two directions. Four moves.

**One: put examples in.** Show it three worked examples of the job, then ask for a fourth. That got its name in the paper that introduced GPT-three, on the twenty-eighth of May, twenty twenty: *in-context learning*. The honest part is a footnote in it.

> These terms are intended to remain agnostic on the question of whether the model learns new tasks from scratch at inference time or simply recognizes patterns seen during training

> — *arXiv 2005.14165 v1 (Brown et al.), footnote 1*

The people who named it declined to say it was learning. Two years later a team led by Sewon Min tested that, on sorting-and-choosing tasks, across twelve models.

> randomly replacing labels in the demonstrations barely hurts performance

> — *arXiv abs page 2202.12837 (Min et al.), abstract*

Give it examples with the wrong answers attached and it still mostly works. What they carry is the shape of the job rather than its content — and nothing durable changes either way. When they go, it goes.

**Two: go and find text, and put that in.** Search something, pick the good parts, place them in front of the question. That is retrieval, and its problem is much older. In July nineteen forty-five an American engineer named Vannevar Bush wrote about how we file things alphabetically and then cannot find them, because minds work by association.

> Selection by association, rather than indexing, may yet be mechanized.

> — *Vannevar Bush, 'As We May Think', The Atlantic, July 1945, section 6; https://www.theatlantic.com/magazine/archive/1945/07/as-we-may-think/303881/*

And in the next breath he describes the machine — a desk you store everything in, consulted at speed, and this is his phrase for it: an enlarged intimate supplement to his memory. Finding and remembering, one device, in an essay with no computers in it. He invented none of the technique; he named the problem, eighty-one years ago.

**Three: let it ask for something.** This is the one almost everybody has backwards, so here is the vendor's own sentence from the day it shipped.

> have the model intelligently choose to output a JSON object containing arguments to call those functions

> — *OpenAI, 'Function calling and other API updates', 13 June 2023, section 'Function calling'; Internet Archive capture 20230615023709 of https://openai.com/blog/function-calling-and-other-api-updates*

The model *outputs an object*. It does not place the call. It writes down what it would like done; software reads that, decides whether it is allowed, does it, and types the result back into the conversation as more text. The model never leaves the room — and the worked example on that page labels its middle step as running on somebody else's machine.

**Four: keep something from last time, and put it back.** That is memory — and here is one company's own documentation on its own memory feature.

> Claude only requests memory operations. Your application executes each request against storage you control

> — *Anthropic, 'Memory tool', Claude Platform documentation, section 'How it works'; https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool, read 10 September 2026*

Requests. Executes. Two verbs, two parties, and never the same party.

Examples, found text, a result that came back, a note kept from before: four names for one operation, in two directions.

Two of the four were named six days apart. The twenty-second of May, twenty twenty: the retrieval paper. The twenty-eighth: the one that named in-context learning. And the retrieval paper is not quite the thing that inherited its name — it describes, in its own words, a fine-tuning recipe, the finding half and the writing half trained together.

Then December twenty twenty-one, and WebGPT, which let a model use a web browser. It had ten things it was permitted to write — search this, click that, quote this part — and a rule about everything else.

> If a model generates any other text, it is considered to be an invalid action.

> — *arXiv 2112.09332 v1 (Nakano et al.), Table 1 caption*

And on memory, a parenthesis from the same paper: each step began from a fresh context, so, in their words, the only memory of previous steps is what is recorded in the summary. Memory, in twenty twenty-one, is already a note somebody else keeps.

June twenty twenty-three: OpenAI makes that an ordinary product feature. November twenty twenty-four: the Model Context Protocol, routinely described as giving models new powers — though its own announcement says the problem it solves is that every new data source needs its own custom implementation. It is a standard for plumbing. It does not choose an action, and it does not authorise one.

Now the part that cuts against everything I have just told you.

In January a large group led from Stanford and the Laude Institute compared six products across sixteen models — thirty-two thousand runs, with error bars. Sixteen of those models ran through more than one product. The gaps go from three-tenths of a point to nearly seventeen — and they go both ways. One Google model did better in a plain scaffold the researchers wrote than in Google's own product; one OpenAI model did better in OpenAI's.

So the effect is real. Now the correction, and it is the most useful thing I read all week. A team from University College London, Nanjing University and Tencent built their own terminal benchmark from real recorded sessions. One model, four applications, a spread of seventeen and a half points — then a second column: the same figures, minus the tasks where the plumbing crashed before the model got a turn. There the spread falls to seven and a half. Most of the gap was not the model doing better. It was runs that never started.

What did not shrink was the bill: across those four, cost per solved task ran eight to one.

> Agent frameworks mainly affect cost-effectiveness rather than the underlying capability ceiling

> — *Zhaoyang Chu and colleagues, 'TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks', arXiv 2605.22535 v1, 21 May 2026, section 1 'Introduction', key finding 2*

One last thing, and it is ours. On day one I passed on a rule: treat differences under three points with scepticism until you know the setups matched. Of those sixteen gaps, five are under three. The rule cuts our own evidence, and I am not going to skip it.

Two things, each on a page, on a date.

One. By May twenty twenty-four LexisNexis had a post up promising that its linked legal citations were free of invention — and, in the same paragraph, that no tool of this kind delivers a hundred per cent accuracy. Both sentences together. Then the page changed: between the twenty-first of June and the tenth of August, twenty twenty-five, the title lost the promise, and so did the sentence beneath it. Free of invention became verified and reliable. The address still spells out the old phrase, which is how you find it — and the sentence that never changed is the modest one. Why it was edited, I will not tell you; the captures are dated and you can walk them yourself.

Two. Nine days ago, on the first of September, Anthropic's release notes recorded something small that is exactly today's idea. On the two models named in that entry you can no longer force a tool call: the setting that meant *you must use a tool* now returns an error, leaving *choose for yourself* and *no tools*. What is withdrawn is the ability to make the model ask.

One idea. **A model reads text and writes text.** Everything else is software: choosing what goes in front of it, deciding what to do with what comes back.

And the honest half, which is not what I expected. On the evidence I could find, that software changes what the work *costs* far more reliably than whether it *succeeds*. Sometimes it moves the score enormously; sometimes most of the movement turns out to be the plumbing falling over. If anybody tells you they know the split, ask for the experiment.

Two doors I am opening and not walking through. The same route that brings in useful text can bring in text written by somebody who is not you, and what protects you is not the model being clever — it is the software refusing to act on it. That is a whole episode, and it is coming. And a single tool call is not, by itself, an agent.

Tomorrow we go backwards. All of this has been about the text somebody puts in front of it. Tomorrow: the text it was built out of — where that came from, and who chose.

That was day six. Thank you for listening.

---

## Sources (20)

- Lewis, Perez, Piktus et al., *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks* — 22 May 2020 (v1)
- Brown et al., *Language Models are Few-Shot Learners* — 28 May 2020 (v1)
- Min, Lyu, Holtzman, Artetxe, Lewis, Hajishirzi, Zettlemoyer, *Rethinking the Role of Demonstrations* — 25 Feb 2022 (v1)
- Vannevar Bush, *As We May Think*, *The Atlantic* — July 1945
- J. C. R. Licklider, *Man-Computer Symbiosis*, IRE Transactions on Human Factors in Electronics HFE-1 — March 1960
- Nakano, Hilton et al., *WebGPT* — 17 Dec 2021 (v1)
- Yao et al., *ReAct* — 6 Oct 2022 (v1)
- Schick et al., *Toolformer* — 9 Feb 2023 (v1)
- OpenAI, *Function calling and other API updates* — 13 June 2023
- Anthropic, *Introducing the Model Context Protocol* — 25 Nov 2024
- Anthropic, tool-use and memory documentation (`memory_20250818`) — read 10 Sep 2026
- Anthropic, Claude platform release notes — entry of 1 Sep 2026, read 10 Sep 2026
- Anthropic (Gian Segato), *Quantifying infrastructure noise in agentic coding evals* — 5 Feb 2026
- Magesh, Surani, Dahl, Suzgun, Manning, Ho, *Hallucination-Free?* — 30 May 2024 (v1); 23 Apr 2025 online
- Merrill, Shaw, Carlini et al., *Terminal-Bench* — 17 Jan 2026 (v1)
- Terminal-Bench 2.1 leaderboard — read 10 Sep 2026
- Chu, Hu, Jiang, O'Hearn, Barr, Harman, Sarro, Ye et al., *TerminalWorld* — 21 May 2026 (v1)
- Vats and Golev, *The Scaffold Effect in Coding Agents* — submitted 8 June 2026 (v1)
- Holistic Agent Leaderboard, CORE-Bench Hard — read live 10 Sep 2026
- LexisNexis, *How Lexis+ AI Delivers Trustworthy Linked Legal Citations* — live page + 10 archive captures

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
