# Understanding AI in a Month 1: The Model Is Not the Product

Understanding AI in a Month — Day 1 · 2026-09-02

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

On the twenty-ninth of July, twenty twenty-six, two OpenAI engineers — Ilan Bigio and Ted Sanders — published a short note about a benchmark. Their model had been scoring badly on a test called ARC-AGI-three, so they rebuilt the software around it and ran it again.

Same trained model. Same reasoning tier. Nothing retrained. At the highest of those tiers the score went from thirteen point three to thirty-eight point three, and the model produced about a sixth as much text getting there.

Nothing inside the model changed. Nearly three times the score.

There are two obvious ways to read that, and both are incomplete.

The first: so benchmarks are rigged. If a score nearly triples because somebody changed some software, the number never meant anything. Tempting — and it throws away the thing that did get measured.

The second is the opposite: so the model was better than we thought and the test was unfair to it. Also wrong, and wrong in a way you can check.

Both assume there is one thing being tested and the number belongs to it. Almost nothing you will read about artificial intelligence is a measurement of one thing.

So before the story, the idea.

A trained model is something you call. You hand it text; it hands back text; and that call is over. It keeps what it learned in training, and whatever is in front of it right now. What it does not keep is you: it has no record of the last time you asked it anything unless somebody puts that record back in front of it. It cannot open a file, run a program, or decide to keep going.

So if that is all a model does, how does software that writes code for an hour exist — reading your files, running your tests, fixing what broke?

Somebody wrote ordinary software around it. That software is the harness.

Picture a specialist in a sealed room, who remembers everything they trained for and nothing about your case unless the file goes back through the door. Somebody outside decides what goes in that file and what is left out, because it only holds so much paper. Somebody carries out what the specialist asks for, brings back the result, and decides when the job is done.

None of that is the specialist, and all of it changes what the specialist can accomplish. That is the harness — and it is not a container the clever part sits inside. It decides what the model sees, what it may touch, what it remembers and when it stops.

Now a second word. Scaffolding is the pieces you added because the model could not do something reliably alone — a planner, a summariser, something that checks the work. Scaffolding is compensation, which is why it is the part that can be taken out later.

Put those together and you get the third word. The system is the whole arrangement that does the job: the model, the harness, the tools it can reach, the limits it runs under, and the thing that scores it. The product you use is a system with a name on it.

Here is the part I would keep if you kept nothing else. Where you draw the line between the model and the system is a choice you make for a question, not a fact you discover. And whichever line you draw, the score belongs to everything inside it.

That is not a new idea, and it is not an idea about artificial intelligence. In nineteen eighty-eight a handful of workstation vendors founded the Standard Performance Evaluation Corporation, recognising, in its own account, a desperate need for realistic, standardised performance tests. Their rules are still published, and rule one point two point two is titled Conditions of Observation.

> SPEC therefore requires that a published result include a description of all performance-relevant conditions.

> — *Standard Performance Evaluation Corporation, 'SPEC CPU 2017 Run and Reporting Rules', rule 1.2.2 'Conditions of Observation'; https://www.spec.org/cpu2017/Docs/runrules.html, read 2 September 2026*

Same chip, different compiler, different score. They concluded a result comes with a configuration sheet. Nearly forty years on, benchmarks for AI systems have inherited that problem and are behind on the paperwork.

ARC-AGI-three is a set of small video games, and a system has to work out the rules by playing them. ARC Prize, who built it, supply a runner that is deliberately thin: every provider gets the same one, so scores compare across companies.

Two things about that thinness mattered. Here is OpenAI on the first.

> First, we noticed that after each game action, all private reasoning was discarded.

> — *OpenAI (Ilan Bigio, Ted Sanders), 'How enabling two settings tripled our scores on the ARC-AGI-3 benchmark', 29 July 2026, section 'ARC-AGI-3', the paragraph beginning 'First, we noticed'; https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/*

They are careful about what survived: the record of past moves stayed; the thinking behind them did not. And the second thing.

> Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew.

> — *OpenAI (Ilan Bigio, Ted Sanders), 'How enabling two settings tripled our scores on the ARC-AGI-3 benchmark', 29 July 2026, section 'ARC-AGI-3', the paragraph beginning 'Second, we saw'; https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/*

So the specialist's private working-out was binned after every move, and the oldest pages of the file were sliding off the desk.

What OpenAI changed is more than the headline suggests. They did not flip two switches inside somebody else's runner. They rebuilt it against their own interface, which moved the memory off the machine running the test and onto their servers; they kept the reasoning between turns; they summarised the record when it filled rather than dropping the oldest part; and they moved the cut-off from a hundred and seventy-five thousand characters to the same number of tokens — which OpenAI describe as coming out about the same. That is their assessment of their own experiment, not an outside one.

Thirteen point three, to thirty-eight point three, at their highest reasoning tier.

Three things about those numbers. First, where they come from. The thirteen point three is not only OpenAI's word for it — ARC Prize publish thirteen point three three for the same model on the same set, measured by them. The thirty-eight has one publisher, and no independent reproduction has been published.

Second, which test it is. Both are on what ARC Prize call the public demonstration set, and their technical report is direct about that.

> Because it is impossible to ensure that system designers don't use the public environments as part of their work, and because the public set is materially easier than the private set, we will never report public set scores of any system on the official leaderboard.

> — *ARC Prize Foundation, 'ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence', arXiv:2603.24621, v2 of 17 April 2026, section 3, on the public demonstration set; https://arxiv.org/html/2603.24621v2, read 2 September 2026*

> As printed in the source: “Because it is impossible to ensure that system designers don’t use the public environments as part of their work, and because the public set is materially easier than the private set, we will never report public set scores of any system on the official leaderboard.”

That is the test's own authors, about the numbers in this story. If you saw this reported as one company's model beating another's, that comparison ran between two different datasets.

Third, what the number is. It is not the fraction of puzzles solved. It is an efficiency score: how many moves the system took against a benchmark of good human play, squared, then capped by how many levels it finished. Thirty-eight is not thirty-eight percent right.

ARC Prize replied publicly the next morning, the thirtieth of July. They called it a real and useful result — about harness design, which is not the same as accepting the score. The thin runner is deliberate, they said: everyone gets the same one, so nobody can quietly tune the scaffolding to the test.

Four hours later François Chollet, who created ARC, posted his own line: a harness built specially for the benchmark is not allowed, general-purpose settings available to every customer are fine. That leaves providers tested under different settings, and he drew the line there.

> My take is that this is fine as long as the settings and the cost are clearly reported.

> — *Francois Chollet (@fchollet), post on X, 30 July 2026 07:37:08 UTC, final paragraph of the post; https://x.com/fchollet/status/2082732210436575669, read 2 September 2026*

Now the part that stops this collapsing into more scaffolding is better, and it comes from ARC Prize's own site. They run a second, community leaderboard for exactly these harness-driven results, on the same twenty-five games and the same measure. Three teams reached ninety-nine, ninety-nine point nine and a hundred percent there in July — the last of them on the same day as OpenAI's note. And ARC Prize's own entry — a full agent harness, with memory and the ability to run code — scores five point two, well below what the thin runner manages on its own.

Hold that against a second posture. In March, Anthropic published a note on designing harnesses for software that runs for hours. When a stronger model arrived, the engineer who wrote it, Prithvi Rajasekaran, removed one piece of scaffolding that had stopped earning its keep and kept two others. He ends not on a result but on a conviction, and says so.

> From this work, my conviction is that the space of interesting harness combinations doesn't shrink as models improve.

> — *Anthropic (Prithvi Rajasekaran, Labs), 'Harness design for long-running application development', 24 March 2026, closing paragraph; https://www.anthropic.com/engineering/harness-design-long-running-apps, read 2 September 2026*

One more measurement, from outside this argument. In February, Anthropic's Gian Segato held the model, the harness and the tasks fixed and varied only the machine resources. Across the ordinary middle of that range the score moved under two points; at the extremes, six. The recommendation: treat leaderboard gaps under about three points with scepticism until you know the setups matched.

Three things, each with a place and a date.

One. As of the first of September, ARC Prize's leaderboard data was unchanged from eleven days earlier — twenty-seven rows, a cost column, no settings column. Chollet's condition was settings and cost clearly reported, and the cost column already exists. Watch for a settings column.

Two. The second and final ARC-AGI-three milestone prize closes on the thirtieth of September, and that track runs with no internet access, so no commercial service can enter. On the twenty-fourth of August the high score was four point five eight percent, and ARC Prize said it got there because one team open-sourced their solution and others built on top of it. By the thirty-first it stood at seven point five one, from a different entrant — announced with no explanation at all, and nothing said about what it ran on. Watch whether the winner clears seven point five one.

Three. Watch what the next harness note from any lab adds, rather than only what it takes away. Since March no frontier lab has published one. An academic team did, on the twenty-eighth of August, and every part of it is an addition.

One idea.

Any benchmark where the system has to act, more than once, and remember, is partly a test of the software around the model. Not because the test is rigged. Because the line between the two is drawn by whoever is measuring, and the number belongs to everything inside it.

So when you read that something scored some number, two questions cost you nothing. What exactly was scored — which test, which version, at which setting? And what changed? The trained model can be identical while the system around it is not.

Today's words were model, harness, scaffolding, and system. Tomorrow we go inside the model, and ask what actually reaches it when you type a sentence. It explains a failure you have almost certainly seen.

That was day one. Thank you for listening.

---

## Sources (16)

- OpenAI (Ilan Bigio, Ted Sanders), *How enabling two settings tripled our scores on the ARC-AGI-3 benchmark*, and the ten datapoints embedded in its chart — 29 Jul 2026; chart data extracted 2 Sep 2026
- ARC Prize Foundation, *ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence* — arXiv v1 24 Mar 2026, **v2 17 Apr 2026**
- ARC Prize (@arcprize), public reply on X — 30 Jul 2026, 03:37 UTC
- François Chollet (@fchollet), post on X — 30 Jul 2026, 07:37 UTC
- ARC Prize, GPT-5.6 Sol scorecard: 13.33% public, 7.78% semi-private, per-environment table — read 2 Sep 2026
- ARC Prize, community leaderboard — read 2 Sep 2026
- ARC Prize, verified leaderboard data, 27 rows — generated 1 Sep 2026, 20:47 UTC
- ARC Prize, *ARC Prize 2026: ARC-AGI-3 Milestone Prize #1* — 6 Jul 2026
- ARC Prize, *Verified Testing Policy*; and the 2026 competition key dates and Kaggle conditions — both read 2 Sep 2026
- Anthropic (Prithvi Rajasekaran), *Harness design for long-running application development* — 24 Mar 2026
- Anthropic (Gian Segato), *Quantifying infrastructure noise in agentic coding evals* — 5 Feb 2026
- Standard Performance Evaluation Corporation, *SPEC CPU 2017 Run and Reporting Rules*, rule 1.2.2; and SPEC's own account of its 1988 founding — read 2 Sep 2026
- Text REtrieval Conference overview, NIST — read 2 Sep 2026
- Yao, Tan, Liu, Li, Wang et al., *Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows* — 27 May 2026
- Vats and Golev, *The Scaffold Effect in Coding Agents* — 8 Jun 2026
- openJiuwen Team et al., *openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents* — 28 Aug 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
