# Understanding AI in a Month 7: Where Training Text Comes From

Understanding AI in a Month — Day 7 · 2026-09-11

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it.*

---

On the second of August this year, a European regulator got the power to fine an AI company that had not published a summary of what its model was trained on.

For new models, the duty to publish had applied for a year already. In January, three researchers went looking for the summaries it required. They ran what they call an exhaustive search, and found five. Four came from an open-source project, a Swiss university consortium, a Polish collaboration and a small image company. The fifth was a file in a model repository that they could not confirm was a summary at all.

Then came the big ones, their summaries dated in the six weeks before that second of August. OpenAI. Google. xAI. Anthropic. The same team's tracker now lists forty-four. I cannot show you that one caused the other. The dates are the dates.

They are forms, with tick-boxes. Where OpenAI's asks it to list the large public datasets it used, the answer is one sentence long.

> The training data for GPT-5.5 includes text from Common Crawl.

> — *OpenAI, 'Public Summary of Training Content for GPT-5.5', version of the summary v1, last update 29 July 2026, section 2.1, field 'List of large publicly available datasets'; cdn.openai.com/pdf/gpt-5-5-eu-ai-act-public-summary-of-training-content.pdf, read 1 September 2026*

That is the whole answer in that box.

Two readings, and I want to head off both.

The first: at last, we know what is in there. There is a line headed *summary of the most relevant domain names crawled*. I have read six of these summaries, and on that line not one names a single website. OpenAI and xAI give nearly the same list of categories, in nearly the same order. Anthropic's answer names the endings: dot com, dot org, dot net.

The second reading is the opposite: it is all theatre, they hoovered up the internet, nobody chose anything.

Wrong too, and today is about why. Somebody always chose. What changed is *what* they were choosing.

Yesterday was the text software puts in front of a model while it runs. Today goes one level down, to the text it was built out of. On day two, the vocabulary turned out to be a fossil of the pile it was counted over. This is the pile.

Start in nineteen sixty-one, with a shelf of paper.

A corpus just means a body of text collected on purpose. A classic early one is the Brown Corpus: a million words of American prose, all of it published in that single calendar year.

Somebody had to decide what a million words of English *is*. Six people did, at a conference at Brown in February nineteen sixty-three, and the manual says how.

> These figures were averaged to obtain the preliminary set of figures used.

> — *W. N. Francis and H. Kucera, 'Manual of Information to accompany A Standard Corpus of Present-Day Edited American English, for use with Digital Computers', Brown University, 1964, revised 1971, revised and amplified 1979, section 1, Contents; read in the Internet Archive capture web/20240313214707 of icame.uib.no/brown/bcm.html on 1 September 2026*

They averaged their opinions — how much newspaper, how much fiction, how much religious writing counted as English. Verse was out; drama was out. Then, *inside* those categories, samples were drawn at random. The randomness is real, and it runs inside a frame six people drew first.

One more line from that manual.

> For all copyrighted material used, the permission of the copyright holder has been obtained.

> — *Same document as above, Brown Corpus Manual, section 1, immediately before section 2 'Versions of the Corpus'*

Sample A-oh-seven is the New York Times, and the manual says the permission details are listed sample by sample.

Now scale that up ten-millionfold, to the ten trillion tokens on those forms, and something gives. Asking permission item by item stops being practical, and so does reading it. So the balance tips. **At that scale, much of the choosing moves from texts to rules — and the rules do the choosing.**

February twenty nineteen, the paper behind GPT-two. How did they choose what to read?

> we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma

> — *Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, 'Language Models are Unsupervised Multitask Learners' (OpenAI; not on arXiv), section 2.1; cdn.openai.com, read 1 September 2026*

A score of three — and what they take is the page the link points to. They are honest about what that measures.

> This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny

> — *Same document as above, section 2.1, the sentence immediately following the karma rule*

Or just funny. Eight months later a Google team published a corpus with six cleaning rules. Keep only lines ending in a full stop, an exclamation mark, a question mark or a closing quote. Throw away any page containing a word from a published list of dirty, naughty and obscene words. And this:

> Since the curly bracket “{” appears in many programming languages (such as Javascript, widely used on the web) but not in natural text, we removed any pages that contained a curly bracket.

> — *Colin Raffel et al., 'Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer', arXiv 1910.10683 version 1 (23 October 2019), section 2.2, fifth cleaning heuristic*

On that same European form, OpenAI says one signal it uses to decide what *not* to read is the United States Trade Representative's list of notorious markets for counterfeiting and piracy. A list compiled for trade policy, helping decide what a model has read.

Other people went and measured what rules like that did.

In April twenty twenty-one a team at the Allen Institute and the University of Washington documented that Google corpus from outside, as people who had not built it. They published it with the dirty-words list and without; comparing the two, that one list had taken out about a fifth of the words — my arithmetic, on their table. Not evenly: documents labelled African American English were removed at forty-two per cent, against about six for the white-aligned category. Be careful with that, because the authors are. Nobody looked at who wrote anything; they ran a dialect classifier, trained on geolocated tweets, over the documents. Their sentence is that the findings *suggest* the list disproportionately removes documents *detected to be* in those dialects.

Three months later another team went looking for repetition, and found one sixty-one-word passage in that same corpus sixty-one thousand and thirty-six times. It reads like filler off a wedding website; here is how it ends.

> believe me, brilliant ideas would be perfect if it can be applied in real and make the people around you amazed!

> — *Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini, 'Deduplicating Training Data Makes Language Models Better', arXiv 2107.06499 version 1 (14 July 2021), footnote 1, which prints the 61-word C4 sequence in full; the excerpt quoted is its final clause*

The recipe had a rule against repeats: keep one copy of any three sentences that recur. The passage is in there sixty-one thousand times anyway.

Their fix was a far stricter hunt for repeats — repeated passages, near-identical pages. Left to generate freely, models trained on the cleaned text emitted memorised text about ten times less often. It does not solve the problem, and several of the same authors said so the following February.

> However, we find that memorization does still happen, even with just a few duplicates—thus, deduplication will not perfectly prevent leakage.

> — *Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, Chiyuan Zhang, 'Quantifying Memorization Across Neural Language Models', arXiv 2202.07646 version 1 (15 February 2022), section 4.3*

So the fights arrive, and they are not the same fight.

One is in a courtroom and asks whether they were allowed to use it — the settlement a judge in California gave final approval to in July: one and a half billion dollars, about four hundred and eighty-two thousand books. No jury found anyone liable; a judge approved an agreement as fair. A year earlier the same court had ruled that using the books to train was fair use, and that keeping a library of pirated copies was not. The settlement covers books from the pirated collections.

The other fight is in front of a regulator, and asks a different question: will you tell us what it was. That is today's.

Now the same form, filled in by a Swiss university consortium. In the field where OpenAI names one dataset, they name nine, each with a link. The biggest tick-box is *more than ten trillion tokens*. OpenAI, Anthropic and xAI tick it, and so do the Swiss — who wrote the number in as well. And they went back through crawls to two thousand thirteen, removing sites that had opted out of AI crawlers by January twenty twenty-five. Retroactively.

Be fair about why they could. Their own form says: no crawlers of their own, no commercial licensing deals, no user data. My reading: they can list their corpus because it is made only of listable things. Whether that can be done at frontier scale, I have not seen anyone show.

Three things you can check yourself.

One, that tracker: forty-four, against five in January — and it is already behind. OpenAI's summary for its newest model, GPT-six Astra, is dated the second of September — the day before that model went on the European market. Its list of large public datasets: the same single sentence. See whether the tracker catches up.

Two, robots files — the text file a website uses to tell crawlers where not to go. The one Reddit served this morning to an ordinary request is five hundred and thirty-eight bytes and ends *disallow slash*: every crawler asked to stay out, on the site where GPT-two found its links. The New York Times' file had five entries in the archive's capture from August twenty twenty-two. Today it has sixty-nine, and tells sixty of them to stay out entirely — including Common Crawl's.

Three. Yesterday it was a company's own web page, changed between two dated captures. This one is a document the law requires. OpenAI's summary for GPT-five-point-five is live at two addresses, with different text. One says last updated the twenty-sixth of June; the other, the twenty-ninth of July. Both are labelled version one. Both were still downloadable this morning. Why it changed is on no page I can find.

Provenance is four questions: where did this come from, when, what was done to it, and what permission was recorded.

For Brown's million words, the manual records permission sample by sample. In twenty twenty-six, for more than ten trillion, the nearest thing to a list of websites in any of those six summaries is: dot com, dot org, dot net.

Underneath both: **at scale, much of the choosing moves from texts to rules.** That is not nobody choosing. Six people averaging opinions, a score of three on a message board, a punctuation mark, a curly brace, a list compiled for trade policy — each a decision somebody made about what a machine would read.

But it is also not the same as knowing what you got. The fifth that word list removed, and the wedding passage repeated sixty-one thousand times, were found later, by people who went and measured. Somebody chose the rule. The result still had to be checked.

Tomorrow: why more text and more computing made better models — and what a scaling curve can and cannot tell you.

That was day seven. Thank you for listening.

---

## Sources (23)

- Regulation (EU) 2024/1689, Articles 53, 101, 111(3), 113; Commission explanatory notice and template — OJ 12 Jul 2024; template 24 Jul 2025; read 1 and 11 Sep 2026
- California AB 2013 (Stats. 2024, ch. 817), Civil Code §3111; *X.AI LLC v. Bonta*, C.D. Cal. 2:25-cv-12295, Dkt. 35; 9th Cir. No. 26-1591 — approved 28 Sep 2024; order 4 Mar 2026; read 11 Sep 2026
- OpenAI, *Training Data Summary Pursuant to California Civil Code Section 3111* — Internet Archive capture, 21 Jan 2026
- Blankvoort, Pandit, Gahntz, *Quality Assessment of Public Summary of Training Content …*, and their tracker — arXiv v1, 26 Feb 2026 (FAccT 2026); tracker read 1 and 11 Sep 2026, last modified 31 Aug 2026
- OpenAI, *Public Summary of Training Content* for GPT-5.5 (both live copies), GPT-5.6 Luna and GPT-6 Astra — GPT-5.5 v1, 29 Jul 2026 and v1, 26 Jun 2026, both re-fetched 11 Sep 2026; Luna v1, 23 Jul 2026; GPT-6 Astra v1, 2 Sep 2026, fetched 11 Sep 2026
- Anthropic (Claude Opus 5), Google (Gemini 3 Pro family), xAI (Grok 4.5), Inkling, Meta (Muse Spark) training-content summaries — 23 Jul, 2 Jul, 8 Jul, 15 Jul and 4 Aug 2026, each from the document's own date field
- Swiss AI Initiative, *Apertus EU Public Summary* — V1, 1 Sep 2025
- Francis and Kučera, *Manual of Information … A Standard Corpus of Present-Day Edited American English* — 1964; rev. 1971; rev. and amplified 1979 (text of the 1979 revision)
- Radford, Wu, Child, Luan, Amodei, Sutskever, *Language Models are Unsupervised Multitask Learners* — Feb 2019
- Raffel, Shazeer, Roberts, Lee et al., *Exploring the Limits of Transfer Learning …* — arXiv v1, 23 Oct 2019; v3, 28 Jul 2020; v4, 19 Sep 2023
- Gao, Biderman, Black et al., *The Pile* — arXiv v1, 31 Dec 2020
- Dodge, Sap, Marasović et al., *Documenting the English Colossal Clean Crawled Corpus* — arXiv v1, 18 Apr 2021
- Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini, *Deduplicating Training Data …* — arXiv v1, 14 Jul 2021
- Carlini, Ippolito, Jagielski, Lee, Tramèr, Zhang, *Quantifying Memorization …* — arXiv v1, 15 Feb 2022
- OpenAI, *GPT-4 Technical Report* — arXiv v1, 15 Mar 2023
- Penedo, Kydlíček, Ben allal et al., *The FineWeb Datasets* — arXiv v1, 25 Jun 2024
- *The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text* — arXiv v1, 5 Jun 2025
- Longpre, Mahari, Lee, Lund et al., *Consent in Crisis* — arXiv v1, 20 Jul 2024
- Wan, Klyman, Kapoor, Maslej, Longpre, Xiong, Liang et al., *The 2025 Foundation Model Transparency Index* — 2025 report; finalisation and release Sep–Dec 2025
- *Bartz v. Anthropic PBC*, Dkt. 231 and Dkt. 680 — 23 Jun 2025; 20 Jul 2026
- Statement of Interest of the United States, *In re OpenAI, Inc. Copyright Infringement Litigation*, 25-md-3143 (S.D.N.Y.) — 1 Sep 2026
- Common Crawl overview, errata, and top-500 domains of CC-MAIN-2026-34 — read 1 Sep 2026
- nytimes.com/robots.txt and reddit.com/robots.txt, live and archived — read 1 and 11 Sep 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
