# Understanding AI in a Month 2: How Text Becomes Tokens

Understanding AI in a Month — Day 2 · 2026-09-05

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

Ask one of these systems how many times the letter R appears in the word strawberry, and for a long stretch it would tell you two.

You have probably also heard the explanation, because it is everywhere. The model does not see the word. It sees chunks — S-T-R, A-W, berry — so the R's get scattered and it cannot count them.

Memorable, repeatable — and this morning I ran the software on the actual question.

In the sentence a person really types — how many R's are in strawberry — the word does not arrive in three chunks. It arrives as one. One symbol. That held in all five published vocabularies I ran, from twenty nineteen to this year; and in the biggest of them, Google's Gemma four, it is one symbol even standing alone at the start of a line.

So the popular explanation is wrong about the case it is famous for. And the truth is worse. The word does not turn up awkwardly broken. It turns up as one symbol with nothing visible inside it.

Two ways people usually take that, and both go wrong.

The first: this is a bug, somebody will fix it. It is not. It is the consequence of a decision made before the model was ever trained — and mostly a good decision.

The second is the opposite. A thing that cannot count letters cannot understand anything, so the whole enterprise is hollow. That feels rigorous and it is unfair — like deciding somebody cannot read because they cannot tell you how many serifs were on the page.

Yesterday was about the software around the model. Today is what happens between your keyboard and the model, on every request.

A model does not read characters. Before your text gets anywhere near it, a separate program cuts it into pieces from a fixed list and hands over numbers pointing into that list. The list is decided once, before training, and frozen.

So what should the pieces be? Letters: nothing is ever unrepresentable, but the sequences get very long, and the work grows sharply with length. Whole words: short sequences, but it breaks the first time somebody types a name, a typo, a word from another language — and there is always such a word.

So the systems in wide use landed in the middle: pieces of words. Common words get one piece; rare words get assembled from several.

How do you pick them? Here is the strange part.

The method is called byte pair encoding, and it did not come from language research. It was a data compression algorithm, published by an engineer named Philip Gage in the C Users Journal in February nineteen ninety-four. His opening line is about disk space.

> Data compression is becoming increasingly important as a way to stretch disk space and speed up data transfers.

> — *Philip Gage, 'A New Algorithm for Data Compression', The C Users Journal 12(2), February 1994, the article's opening paragraph; read from the Internet Archive capture of 29 March 2021, https://web.archive.org/web/20210329111703/http://www.pennelynn.com/Documents/CUJ/HTML/94HTML/19940045.HTM*

The algorithm is almost childishly simple. Find the pair of adjacent bytes that occurs most often in your data. Replace every instance of it with a byte that was not in the original data. Do it again. Stop when there are no more repeated pairs, or when you run out of unused bytes.

Twenty-one years later, Rico Sennrich, Barry Haddow and Alexandra Birch pointed that trick at machine translation. Their problem was rare words: a translation system has a fixed vocabulary and language does not.

And here is the adaptation the popular telling drops, which is the one that matters. Gage stops when he runs out of unused bytes, because he is compressing a file and every symbol has to be a real byte. If you are building a vocabulary instead, a new symbol is just a new entry on a list. The ceiling comes off. Gage's version could never hold more than two hundred and fifty-six symbols. The encodings OpenAI publishes today hold two hundred thousand.

And Sennrich and his colleagues claimed far less than the legend does. Their own summary is nought point three to one point three BLEU over a dictionary fallback — and on English to German the best number in their table is not the byte-pair system at all. It is a character-bigram one.

So: run the merges over an enormous pile of text, stop after a chosen number of passes, and freeze. That list is the vocabulary. The rules for cutting are designed; which words come out cheap is a fossil of the pile.

One more piece. The list also holds entries you cannot type. In OpenAI's published chat format, the marker that opens a turn is symbol number two hundred thousand and six — while assistant is just an ordinary word among the ordinary ones. There is no field for who is speaking. A conversation is one flat run of symbols, and what separates the application's frame from what you wrote is that your keyboard cannot reach those numbers.

So watch the vocabularies grow. Twenty nineteen: fifty thousand two hundred and fifty-seven. Twenty twenty-two: a hundred thousand. May twenty twenty-four: two hundred thousand. This year, Google's Gemma four: two hundred and sixty-two thousand.

And here is what that does not show. Standing alone at the start of a line, strawberry is still three pieces in four of the five I ran — including the second biggest. Only Gemma four keeps it whole, and Gemma differs from the others in scheme as well as in size. What makes the word one symbol in the sentence people actually type is smaller and stranger: the space in front of it belongs to the symbol.

In February twenty twenty-four Andrej Karpathy published a lecture called Let's build the GPT Tokenizer, and wrote it up as a text chapter still in his repository today. Where he motivates it, he says:

> Tokenization is at the heart of a lot of weirdness in L L M s and I would advise that you do not brush it off.

> — *Andrej Karpathy, 'LLM Tokenization', lecture.md in his minbpe repository on GitHub, the written lecture accompanying the video 'Let's build the GPT Tokenizer' (February 2024), section 'Brief taste of the complexities of tokenization'; read 5 September 2026*

> As printed in the source: “Tokenization is at the heart of a lot of weirdness in LLMs and I would advise that you do not brush it off.”

And now the careful part, because it would be easy to let that eat everything. When Xiang Zhang and two colleagues measured counting in twenty twenty-four, every mistake under ordinary tokenisation was an under-count, never an over-count — exactly what you would expect if the letters inside a symbol are invisible.

But a benchmark called CharBench, published this year by Omri Uzan and Yuval Pinter, splits the claim in two. They found tokenisation only weakly correlated with counting correctly — and strongly implicated in locating a letter inside a word, where accuracy falls as the token containing it grows longer. The vocabulary explains position better than it explains the count.

Two things happened next, and they pull against each other.

First, the fairness problem got numbers — and then a price tag.

If the vocabulary is a fossil of a pile of text, the cost of writing in your language depends on whose text was in the pile. I measured that this morning, over a hundred sentences in forty-one languages. In the twenty nineteen vocabulary, Burmese cost fifteen and a half times what the same sentences cost in English; by the twenty twenty-four one, three times. Tamil went from fourteen and a half to about two. An enormous improvement, and it deserves to be said first.

And then Lao. Twelve times in twenty nineteen; still seven and a half times in twenty twenty-four. The gap closed most for the languages already closest, and least for the ones furthest away.

And it is not an abstraction. Anthropic's published pricing page says today that its models from four-point-seven onward use a newer tokeniser producing approximately thirty per cent more tokens for the same text — and its published Opus price is five dollars per million input tokens on both sides of that change. Same price per token, about a third more tokens. Which is why I would file a tokeniser change and a price change as the same kind of event.

The second is a research programme trying to delete this step altogether — byte-level models that learn where the boundaries go instead of being handed them. You can download one.

But here is the sharpest objection to my own framing. Last September Catherine Arnett wrote There is no such thing as a tokeniser-free lunch. Her line is this.

> No matter how you chunk up your input data, you’re doing tokenization.

> — *Catherine Arnett, 'There is no such thing as a tokenizer-free lunch', Hugging Face community article, 25 September 2025, section 'What is Tokenization and Why Does it Exist?'; https://huggingface.co/blog/catherinearnett/in-defense-of-tokenizers, read 4 September 2026*

And the files prove her point. The Allen Institute's byte-level model ships a tokeniser class of its own, and a vocabulary of five hundred and twenty symbols — down from the hundred thousand it was converted from, which the same file still records. The vocabulary did not vanish. The argument is about how coarse the pieces are and who picks them, not whether there are pieces.

Two things, both checkable rather than atmospheric.

One. OpenAI publishes the library that maps its model names to vocabularies. On the seventeenth of August this year one line of it was widened: the entry used to read gee-pee-tee-five-dash, and now the dash is gone, so any name starting with those five characters gets the vocabulary added in May twenty twenty-four. The commit is titled Fail open for GPT-five model tokenisers, and the file's own comment says prefix matching can match models that do not exist. So it is a claim by a library, not a fact about a server — and the day a name in it points somewhere new, the list was rebuilt.

Two. When a lab releases a model, check whether its download page carries a tokeniser file, and what size it claims. I checked eight this morning, three of them released in the last week: every one has a file, from a hundred and twenty-nine thousand entries up to two hundred and sixty-two thousand. A byte-level model answers with three digits. If a flagship ever does, that is the day this step stopped being universal.

One idea.

The model does not see what you typed. It sees a run of symbols from a list frozen before it was trained, chosen by a compression algorithm run over a pile of text nobody showed you.

Karpathy's own line is that a lot of what looks like a problem with the neural network traces back to this instead. So does a fairness question that looks like an engineering detail and is not.

And hold on to how thin the ice was under the version you already knew. Str, aw, berry was a good story, told confidently, for two years. It was one measurement away from being wrong.

Tomorrow: given a run of those symbols, what does the model actually do? One operation, repeated — and simpler than almost anyone expects.

That was day two. Thank you for listening.

---

## Sources (14)

- Philip Gage, *A New Algorithm for Data Compression*, *C Users Journal* 12(2), pp. 23–38 — February 1994
- Rico Sennrich, Barry Haddow, Alexandra Birch, *Neural Machine Translation of Rare Words with Subword Units* — arXiv v1 31 Aug 2015
- Andrej Karpathy, *Let's build the GPT Tokenizer*, and the written lecture in `karpathy/minbpe` — video 20 Feb 2024; lecture re-fetched 5 Sep 2026, sha256 `1dcf4236f8c8…`, byte-identical to the previous day's copy
- OpenAI, `tiktoken`: `tiktoken_ext/openai_public.py`, `tiktoken/model.py`, `CHANGELOG.md` — `o200k_base` added 13 May 2024 (commit `9d01e56`, v0.7.0); language comment added 2 Oct 2024 (commit `05e66e8`); prefix widened 17 Aug 2026 (commit `212b893`); all read 5 Sep 2026
- Google, Gemma 4 12B IT `tokenizer.json` (32,169,626 bytes, sha256 `cc8d3a0c…`) and `tokenizer_config.json` — re-fetched 5 Sep 2026, byte-identical to the 4 Sep copy
- Published configurations: DeepSeek V4 Pro, GLM-5.3, Kimi K3, Qwen3.8-27B, `gpt-oss-120b`, Gemma 4 — `config.json` and `tokenizer_config.json` re-fetched 5 Sep 2026
- FLORES-101 development split, 41 languages × 100 parallel sentences — re-encoded 5 Sep 2026; every relative figure identical to the 4 Sep run
- Petrov, La Malfa, Torr & Bibi, *Language Model Tokenizers Introduce Unfairness Between Languages* — arXiv v1 17 May 2023
- Zhang, Cao & You, *Counting Ability of Large Language Models and Impact of Tokenization* — arXiv v1 25 Oct 2024
- Omri Uzan & Yuval Pinter, *CharBench: Evaluating the Role of Tokenization in Character-Level Tasks* — arXiv 2508.02591, v3 6 Apr 2026; AAAI-26
- Pagnoni et al., *Byte Latent Transformer*; Hwang, Wang & Gu, *Dynamic Chunking*; Minixhofer et al., *Bolmo*; Kallini, Pagnoni et al., *Fast Byte Latent Transformer* — 13 Dec 2024; 10 Jul 2025; 17 Dec 2025; 8 May 2026
- Catherine Arnett, *There is no such thing as a tokenizer-free lunch* — 25 Sep 2025
- Anthropic developer pricing documentation: the tokeniser note and the published per-token prices — read 5 Sep 2026
- `allenai/Bolmo-7B` `config.json`, `tokenizer_config.json` and the HuggingFace model API download counts — read 5 Sep 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
