# Understanding AI in a Month 11: What Does a Score Prove?

Understanding AI in a Month — Day 11 · 2026-09-19

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it.*

---

On the third of September, OpenAI launched GPT-six Astra. It introduced the model as

> the world’s most intelligent and aligned model.

> — *OpenAI, Introducing GPT-6 Astra (launched 3 September 2026), opening sentence; archive.today capture 20260907123827 of openai.com/index/gpt-6-astra/*

Below that came the results, each set beside rival models. Beneath the results, a single line.

> Evaluation scores are the maximum at any effort.

> — *OpenAI, Introducing GPT-6 Astra, note beneath the results tables, before the footnotes; archive.today capture 20260907123827*

And one row of those results was labelled as coming from outside OpenAI: an overall index run by an independent firm, Artificial Analysis. On that row, as OpenAI printed it, Astra scored sixty-one point two. Three models from Anthropic scored higher. The highest was sixty-five point seven.

All of that is on one page. Yesterday was a score used to train a model. Today is a score used to measure one — and what, exactly, it proves.

A launch like that invites two readings, and both go too far.

The first is that the highest number is the verdict. Most intelligent is a claim about a quality. Each row measures something much narrower. And the two rows built as overall indexes, both labelled as Artificial Analysis's, didn't put Astra first as OpenAI printed them.

The second reading is the cynical one: launch charts are marketing, so ignore them. That throws away real information. Take one row, a test of work at a computer terminal. OpenAI's figure for Astra was fifty-seven point nine per cent. Artificial Analysis ran the test itself, every task three times, and got fifty-nine point one — a little higher, with the same order for the four models it tested. That number survived the check.

So some numbers hold up and some mean less than they seem to. The skill is telling which is which.

The idea you need is older than computers doing any of this. It's called construct validity, and it comes from psychology.

In nineteen twenty-three, the psychologist Edwin Boring wrote about intelligence tests in The New Republic. One line from that article is still quoted.

> Intelligence is what the tests test. This is a narrow definition, but it is the only point of departure for a rigorous discussion of the tests.

> — *E. G. Boring, Intelligence as the Tests Test It, The New Republic 35 (6 June 1923) pp. 35-37, section 'What the Tests Test'; Mead Project transcription, read in an Internet Archive capture of 2023*

Read the second sentence and it isn't cynical. It's a starting point. Until you know more, the score is the only thing you've pinned down — and, he wrote, the everyday meaning of intelligence is much broader.

Between nineteen fifty and nineteen fifty-four, a committee of the American Psychological Association worked out what should be checked before a psychological test is published. Its chief new idea was a term, first drafted by a subcommittee that included Paul Meehl. In nineteen fifty-five, Meehl and Lee Cronbach explained it. First, the word construct.

> A construct is some postulated attribute of people, assumed to be reflected in test performance.

> — *L. J. Cronbach and P. E. Meehl, Construct Validity in Psychological Tests, Psychological Bulletin 52(4) 1955, section 'Kinds of Constructs'; Classics in the History of Psychology text, http://psychclassics.yorku.ca/Cronbach/construct.htm*

Postulated means assumed, not seen. Nobody has ever seen intelligence. You see answers, and infer a quality behind them. Construct validity is the case that the answers really do reflect that quality. And they said when you can't skip making the case.

> Construct validity must be investigated whenever no criterion or universe of content is accepted as entirely adequate to define the quality to be measured.

> — *Lee J. Cronbach and Paul E. Meehl, 'Construct Validity in Psychological Tests', Psychological Bulletin 52, 1955, section 'Four Types of Validation', p. 282; Classics in the History of Psychology text, http://psychclassics.yorku.ca/Cronbach/construct.htm*

In plain words: when no test simply is the thing, you have to argue that yours tracks it.

Now swap people for models. The name on a benchmark — reasoning, coding, intelligence — is a construct. The questions are what actually got measured. A score tells you how a model did on those questions. That it measures the name is a second claim, and somebody has to make the case.

How often does anyone? Last November, a team with Andrew Bean of Oxford as first author, and twenty-nine expert reviewers, went through four hundred and forty-five benchmarks for language models, published in peer-reviewed papers at leading research conferences — not the labs' own in-house tests. Just over half gave evidence that they measured what they said they measured. Sixteen per cent used any statistical test or uncertainty estimate when comparing results.

Two more ideas complete the picture. The second is elicitation: how the model was run. Which instructions, how long it may think, which tools, how many tries. Change them and the score moves. OpenAI's own safety framework, in April twenty twenty-five, said this about its tests for dangerous abilities.

> we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling, on capabilities that may emerge in real world use and misuse.

> — *OpenAI, 'Preparedness Framework', Version 2, last updated 15 April 2025, section 3.1 'Evaluation approach', p. 8; https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf*

That was written about safety testing, but by my reading the logic carries to a launch chart. A score is what one setup got out of a model. It isn't everything the model can do, and it isn't what a rival would score set up the same way.

The third idea is design: how the score was counted. How many questions stand behind it? Does one attempt count, or the best of several? Who marked the answers? In twenty twenty-four, Evan Miller at Anthropic argued that evaluations are experiments, and that their results should come with error bars, so a reader can tell a real difference from noise.

So, three questions. What do the questions actually ask? How was each model run? And how was it counted?

Now back to the launch page, with those three questions in hand. Most of the answers are in its footnotes.

How was each model run? Maximum at any effort means that for each test the page shows the best result across effort settings, chosen once the results were in. Two cyber tests were run, OpenAI says, without production safeguards — by my reading, the safety tester's way of pushing for the highest number — and one of them, a perfect score, sits in the page's opening paragraph. One of those was also run without its six-hour time limit; OpenAI says the models are fast enough that it made little difference. And on the puzzle test ARC-AGI-three, Astra ran in OpenAI's own harness, the software between model and test. That was day one's story, and it has a sequel. The ARC Prize Foundation, which runs the test, ran Astra both ways: sixty-two point seven per cent on its own neutral setup, ninety-nine point nine with the setup that uses OpenAI's features.

How was it counted? On one test of reverse-engineering software, OpenAI's text gives eighty-eight per cent of tasks solved in a single attempt, and ninety-nine point two within four attempts. Both are labelled; the page's results grid carries only the single-attempt figure. On a medical test, OpenAI ran the rival models itself, with one of its own models marking the answers, following what it calls the test's intended procedure. And for two rows, the scores in Anthropic's Fable columns came, in OpenAI's words, from Mythos, which is Fable with fewer safeguards.

And what do the questions ask? Astra's page says it saturates ARC-AGI-three. The foundation called Astra meaningful progress, and it described its own test like this.

> its environments have deterministic, closed-ended mechanics and goals. It does not represent the complexity and open-endedness of the real world.

> — *ARC Prize Foundation (Greg Kamradt), OpenAI's GPT-6 Astra on ARC-AGI-3, arcprize.org/blog/astra, published 3 September 2026, section 'ARC-AGI Series'*

That's a test's makers naming the distance between the name and the items.

Then the independent index moved, and it's worth watching how.

The row on OpenAI's page was version four point one point one of Artificial Analysis's index. On the fourth of September, the day after the launch, the firm put out version four point two: Anthropic's Claude Fable five point one leading, followed by Astra. On the seventh came version four point three, which swapped in two newer tests — and Astra and Fable five point one tied, at fifty-three. The firm gave its reasons: an index closer to real-world problem solving, more private test sets to prevent gaming, and less saturation. Nothing in the record says either model changed. The index did: its tests, how much each counted, and how answers were marked.

That doesn't make the index wrong. It makes it a design, like every other score. And it cuts both ways. Anthropic's Fable models can decline a request, and on the firm's runs the answer then comes from another Claude model. The firm labels those runs with fallback, and on its first of September figures, the fallback produced about four per cent of the output tokens.

The foundation's response was different again. It says its leaderboard will now report both harness results, each condition clearly labelled. On the first of September that leaderboard had a column for cost and none for settings.

Three things you can check.

One. The ARC-AGI leaderboard on the foundation's website. The page's key already names both harnesses; check whether each of Astra's rows does, as promised on the third of September. One catch: its default view shows only systems that cost under ten thousand dollars to run, and every Astra run the foundation reported cost more.

Two. Artificial Analysis's changelog. The index has been through three versions since early August, and on the fourth of September the firm said it was already at work on a version five. When that appears, check whether the order at the top moves, and read the reason given for each test added or dropped.

Three. OpenAI's Astra page. As saved on the seventh of September, it still showed version four point one point one of the index. Watch whether that row is updated, and to what.

The idea to keep is construct validity: whether a test measures the quality in its name. Edwin Boring's point still stands. The score is where you start.

So before you believe a benchmark claim, ask three things. What do the questions actually ask, and does that match the name? How was each model run — and were they all run the same way? How was it counted — how many questions, how many tries, and who marked them? Then the question under all three: what would have to be true for this number to mean what the headline says?

To read more, the encyclopedia has articles on construct validity, capability elicitation, and AI evaluations.

Tomorrow: why models think longer, and what that extra thinking buys.

That was day eleven. Thank you for listening.

---

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
