# Understanding AI in a Month 20: Why Benchmark Scores Rot

Understanding AI in a Month — Day 20 · 2026-10-01

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

On the fourth of September, Artificial Analysis, an independent company that publishes rankings of AI models, changed what goes into its headline score. It added two tests and took one out. Its changelog gave the reason in a single line.

> G P Q A Diamond, an exceptional scientific reasoning evaluation that has now been saturated

> — *Artificial Analysis, 'Announcing Artificial Analysis Intelligence Index v4.2', 4 September 2026, changelog list under the heading 'Intelligence Index v4.2 changelog'; https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2*

> As printed in the source: “GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated”

On X, the company added a second reason: the test's multiple-choice format doesn't reflect real tasks in those fields. And a study published at this year's International Conference on Machine Learning looked at sixty widely used tests of language models and found, by its own measure, that twenty-nine had largely lost the power to tell the leading models apart. Today: why a test that once measured something stops measuring what people think it does.

Two readings of news like this mislead.

The first is that saturated means mastered: the machines have learned the subject, so the test has nothing left to ask. The study's authors mean something narrower. The best five models can't be reliably told apart, and they're bunched near the highest score anyone has reached. That needn't be a hundred per cent. One test, LiveBench, came out as very highly saturated with its leaders clustered around seventy-nine per cent, which the authors read as models stalling, not the task being finished. And they add a condition: if a test was valid in the first place, saturation can mean the task is solved. Whether it measured what its name says is a separate question, and a crowded top score doesn't answer it.

The second reading is the cynical one: high scores just mean the models memorised the answers, so benchmark progress is fake. Leakage is real, and we'll get to it. But a test can also stop being useful because the models genuinely got better. Contamination and overfitting are possibilities to investigate, not conclusions you can read off a high score. The useful question isn't whether a score has rotted. It's how, and how much.

A benchmark score is a claim: that a small, fixed set of questions stands in for a skill. There are three common ways that claim decays, and each breaks a different link.

The first is contamination: the test leaks into what the model learned from. A student who has seen the exam paper can score well without knowing the subject. Language models learn from enormous scrapes of the internet, and benchmarks are published on the internet. In May twenty twenty, OpenAI's paper introducing GPT-3 described trying to remove benchmark questions from its training data. Then this.

> Unfortunately, a bug in the filtering caused us to ignore some overlaps, and due to the cost of training it was not feasible to retrain the model.

> — *Tom B. Brown and colleagues (OpenAI), 'Language Models are Few-Shot Learners', arXiv 2005.14165, first posted 28 May 2020 (version 4, 22 July 2020), section 2.2 'Training Dataset', page 9*

So they measured instead, rescoring each test on just the questions they were confident the model hadn't seen. Most scores barely moved, and two got an asterisk. Their own conclusion was careful: either their check had overestimated contamination, or contamination had little effect. In twenty twenty-three, OpenAI's GPT-4 report noted that parts of a test suite called BIG-bench had been mixed into its training data by accident, even though BIG-bench carries a marker string meant to help keep it out. A marker is a label, not a lock.

The second is overfitting: fitting the particular sample instead of the general pattern. A model with enough freedom will find patterns in any sample, including ones that are just accidents of that sample. It learns the noise along with the signal.

With benchmarks, the worry is a whole field doing this slowly, keeping whichever ideas score best on the same public test, year after year. In twenty nineteen, Benjamin Recht and three colleagues at Berkeley tested that worry. They rebuilt the ImageNet image-recognition test from scratch, following the original recipe, and ran the existing models on it. Accuracy fell by eleven to fourteen percentage points. But the models kept almost exactly the same order, and

> accuracy gains on the original test sets translate to larger gains on the new test sets.

> — *Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt and Vaishaal Shankar, 'Do ImageNet Classifiers Generalize to ImageNet?', arXiv 1902.10811, first posted 13 February 2019 (version 2, 12 June 2019; ICML 2019), abstract*

Their results suggested the drop came from the new images being slightly harder, not from years of tuning against the old ones. A study of a hundred and twenty competitions on the Kaggle platform, the same year, found little evidence of substantial overfitting either.

The third is saturation: the test runs out of room. In twenty twenty-one, Douwe Kiela of Facebook AI Research and colleagues put it this way.

> While it used to take decades for machine learning models to surpass estimates of human performance on benchmark tasks, that milestone is now routinely reached within just a few years for newer datasets

> — *Douwe Kiela and colleagues, 'Dynabench: Rethinking Benchmarking in NLP', arXiv 2104.14337, 7 April 2021 (NAACL 2021), section 1 'Introduction', page 1*

One language test, GLUE, saturated within a year, they wrote. A saturated test hasn't necessarily gone wrong. It has stopped separating the systems people want to compare.

Here's one test's life. G P Q A was published in November twenty twenty-three by researchers led from New York University: graduate-level science questions meant to be Google-proof. Its best-checked set, the Diamond questions, kept questions that experts got right and most skilled non-experts, searching the web, got wrong. GPT-4 scored thirty-nine per cent. In the study's table, first posted in February this year, the five best models scored between eighty-three and eighty-eight. This September it left the index. The knowledge test M M L U had run out of room earlier: Humanity's Last Exam was released in January twenty twenty-five because, its authors said, models now scored over ninety per cent on tests like it.

The hardest maths is following. On the tenth of September, the research group Epoch AI posted that every problem in the top tier of its FrontierMath benchmark had now been solved by AI. Epoch counts a problem as solved once any model has solved it in any attempt. Two things belong beside that. Epoch says OpenAI funded the benchmark and has exclusive access to some of its problems, and the model that solved the last one was OpenAI's. And in June, Epoch addressed errors in forty-two per cent of its problems. Near the top, a test's own mistakes start to blur what a higher score means.

The study's wider results carry a caution. Of its sixty benchmarks, twenty-nine showed high or very high saturation. Among those under two years old, about forty-three per cent were saturated; among those over five, about fifty-five. The authors call that trend modest and not statistically significant, so it's a direction, not a law. Bigger test sets went with less saturation, and so, with age muddying the comparison, did questions written by experts.

So what do you do about a rotting test? Two answers are on offer, and they're aimed at different rots.

The first is to keep the questions out of reach. On the fourth of September, Artificial Analysis put forty per cent of its index's weight on private, held-out tests, double the share before, which it says reduces the ability of labs to game evaluations. By my count of its current table, it's now forty-five. LiveBench was designed the other way, replacing a slice of its questions every month so they postdate a model's training.

The second answer comes from the same study, and it's less comforting about the first. Its four private benchmarks saturated much like the fifty-six public ones.

> Hiding test data does not appear to prevent saturation once benchmarks are widely adopted

> — *Mubashara Akhtar, Anka Reuel and colleagues, 'When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation', arXiv 2602.16763, version 4, 6 August 2026 (ICML 2026), section 4.1, paragraph 'Accessibility and task design', page 6*

Four is a small sample, so that's a hint, not a verdict. But the authors also note that even tests built to keep adding new questions can saturate when they can't measure finely enough to separate the leaders, and LiveBench's own changelog describes an update meant to fix saturation and contamination only to some extent. The authors back refreshing too, but their remedies lead with resolution: bigger and harder tests, uncertainty reported beside the score, and rules, written in advance, for revising or retiring a test once it stops separating the leaders.

I think both answers are right, about mostly different diseases. Secrecy and freshness are aimed at contamination, though freshness can slow saturation too. Size, difficulty and honest error bars are what let a test keep separating the leaders. And a private test has a cost of its own: nobody outside can check it. Fresh doesn't mean fair, and private doesn't mean valid.

One. Artificial Analysis's index. The company says version five will raise the private share further and that more interim updates are coming. Through October, watch its changelog: does the private share keep rising past forty-five per cent, which test is retired next, and does it say why?

Two. Whether the live tests stay live. Microsoft's SWE-bench Live promises fifty new Python software issues a month. Check in early November whether October's arrived. LiveBench's public changelog has had no entry since the eighth of January.

Three. Epoch's newer test, FrontierMath Erdős: sixty-eight unsolved problems, launched on the third of September. OpenAI's GPT-6 Astra solved two. Watch how fast that number moves. That's a new test's clock starting.

The idea to keep is that a benchmark score has a shelf life, and three questions tell you how much of it is left. Could the model have seen these questions, or close copies, before the test? Has a whole field been tuning against this one test for years? And can the test still tell the top systems apart, by more than its own margin of error? None of those makes a score worthless. Each tells you how much of its claim is still standing.

To read more: When AI Benchmarks Plateau, by Mubashara Akhtar, Anka Reuel and colleagues, from this year's International Conference on Machine Learning. Its appendix walks through five benchmarks, from saturated to not, in a page.

---

## Sources (23)

- Artificial Analysis, *Announcing Artificial Analysis Intelligence Index v4.2*, artificialanalysis.ai/articles — 4 September 2026
- Artificial Analysis, post on X on Intelligence Index v4.2 — 6 September 2026 (00:28 UTC)
- Artificial Analysis, *Intelligence Benchmarking Methodology*, version 4.3.2 — read 1 October 2026
- Epoch AI, post on X on FrontierMath Tier 4 — 10 September 2026
- Epoch AI, *FrontierMath Tier 4* benchmark page and *Benchmarks* hub — read 1 October 2026
- Akhtar, Reuel, Soni, Ahuja et al., *When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation*, arXiv 2602.16763 v4 (ICML 2026) — 18 February 2026; v4 6 August 2026
- Rein et al., *GPQA: A Graduate-Level Google-Proof Q&A Benchmark*, arXiv 2311.12022 — 20 November 2023
- Brown et al. (OpenAI), *Language Models are Few-Shot Learners*, arXiv 2005.14165 — 28 May 2020
- OpenAI, *GPT-4 Technical Report*, arXiv 2303.08774 — March 2023
- BIG-bench, *training_on_test_set* README (canary GUID), in the BIG-bench repository on GitHub — read 1 October 2026
- Jain et al., *LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code*, arXiv 2403.07974 — 12 March 2024; v2 6 June 2024
- Zhang et al. (Scale AI), *A Careful Examination of Large Language Model Performance on Grade School Arithmetic* (GSM1k), arXiv 2405.00332 — 1 May 2024; revised November 2024
- International AI Safety Report 2026, arXiv 2602.21012 — February 2026; arXiv 24 February 2026
- Recht, Roelofs, Schmidt and Shankar, *Do ImageNet Classifiers Generalize to ImageNet?*, arXiv 1902.10811 (ICML 2019) — 13 February 2019
- Roelofs et al., *A Meta-Analysis of Overfitting in Machine Learning*, NeurIPS 2019 — December 2019
- Kiela et al., *Dynabench: Rethinking Benchmarking in NLP*, arXiv 2104.14337 — 7 April 2021
- Wang et al., arXiv 1804.07461, *GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding* — 20 April 2018
- Wang et al., *SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems*, arXiv 1905.00537 — 2 May 2019
- Microsoft Research, *Microsoft DeBERTa surpasses human performance on the SuperGLUE benchmark* — 6 January 2021
- Gema et al., *Are We Done with MMLU?*, arXiv 2406.04127 (v3) — 6 June 2024; v3 10 January 2025
- Phan et al., *Humanity's Last Exam*, arXiv 2501.14249 (quoted from v11) — 24 January 2025; v11 28 July 2026
- White et al., *LiveBench: A Challenging, Contamination-Limited LLM Benchmark*, arXiv 2406.19314; LiveBench changelog — 27 June 2024; changelog read 1 October 2026
- Microsoft, *SWE-bench-Live* README, in the SWE-bench-Live repository on GitHub — read 1 October 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
