# Understanding AI in a Month 4: How Words Affect Other Words

Understanding AI in a Month — Day 4 · 2026-09-07

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

Here is a sentence. *The suitcase would not fit in the car because it was too big.* Ask what "it" refers to and the natural reading is the suitcase. Change one word at the far end — *because it was too small* — and the natural reading is the car. A word near the end of the sentence settled what a pronoun in the middle referred to.

Getting that to happen inside a machine is what today's operation is for: the step where the pieces of a text are compared against one another and come out changed by it. It is called attention, and the name is the first thing to put down.

The paper that made it the standard design was posted on the twelfth of June, twenty seventeen. This week I put that sentence to two models with the question after it. On a model from twenty nineteen, built on that design, flipping the far word shifted the odds it put on the two answers by about fifteen per cent — and in the wrong direction. On one published this year, by a factor of about three hundred and eighty.

Two readings that will mislead you.

The first is the word. Attention sounds like concentrating, as though something in there were holding a spotlight. The twenty fourteen translation paper that put the mechanism into wide use, by Bahdanau, Cho and Bengio, uses the word three times in the whole paper, all in one paragraph, and the function it names is an alignment model. The spotlight came with the nickname, not the machinery.

The second is more tempting, because it looks like progress. Every permitted pair of words gets a number out of this operation, in every head of every layer — so draw the numbers as a picture and read the reasoning off it. That produced a published argument, and it is beat five.

Two ideas, and the second is built on the first.

The first: **a word becomes a list of numbers, which you can picture as a position in a space.**

That idea predates the machines that use it. Jurafsky and Martin's textbook traces it to three fields in the nineteen fifties; two matter here. From linguistics, the proposition that a word's meaning is the company it keeps — the textbook names Joos in nineteen fifty, Harris in fifty-four, Firth in fifty-seven — and quotes Joos, with an omission in the middle that is the textbook's, not mine.

> the linguist's meaning of a morpheme is by definition the set of conditional probabilities of its occurrence in context with all other morphemes.

> — *Daniel Jurafsky and James H. Martin, 'Speech and Language Processing', 3rd edition draft of 19 August 2026, chapter 5 'Embeddings', Historical Notes, p. 22, quoting Martin Joos (1950); https://web.stanford.edu/~jurafsky/slp3/5.pdf*

> As printed in the source: “the linguist’s “meaning” of a morpheme. . . is by definition the set of conditional probabilities of its occurrence in context with all other morphemes.”

From psychology, seven years later, came the geometry: Osgood and colleagues proposed that a word's meaning could be a point in a space of many dimensions, with similarity as distance. The field's name for that position is an **embedding**. His numbers came from people; a model's are learned from text.

One position per word is not enough, though. A word arrives carrying its average company across everything the model ever read, and needs one reflecting the company it is in now — with a restriction: in a model that writes, a position may use itself and everything before it, and nothing after. The suitcase sentence works because the question comes last.

The operation that updates it has three steps. **Compare, weight, blend.** Learned transformations turn each position's numbers into three versions of itself: one to compare with, one to be compared against, one to be mixed in. The textbook calls those three roles one input plays; the field calls them query, key and value.

Each position compares itself against the ones it may use, itself included, and each comparison gives back a number. Those numbers become weights adding up to one, the others are blended in exactly those proportions, and that blend is then **added to what the position already had** — the part popular accounts drop. Nothing is replaced; information accumulates. The only thing doing spotlight work there is the word attention.

Then it happens again, with its own comparisons learned during training, over what the last round produced. Each pass is a **layer**: the attention step, then an ordinary small network run over each position on its own — and by the twenty seventeen paper's dimensions the second part is the bigger of the two.

One property to carry into beat four: the comparison step has no sense of order. Zero out the numbers that carry position in the twenty nineteen model, shuffle the words before the last one eight different ways, and the attention step's output comes back the same to about one part in seven million. Order is not worked out by the operation. It is added at the door.

The twelfth of June, twenty seventeen. Eight authors, most of them at Google, posted a paper called *Attention Is All You Need*. Three corrections, all from the version posted that day.

It is a machine translation paper, not a chatbot paper. Its headline number, the authors' own measurement, is an English-to-German score of twenty-eight point four on one twenty fourteen test set — more than two BLEU above the best previously reported models, ensembles included.

Second, it invented neither attention nor self-attention: its background section calls self-attention "sometimes called intra-attention" and cites four earlier papers that used it. Here is its claim about itself, hedge and all.

> To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using R N N s or convolution.

> — *Ashish Vaswani and colleagues, 'Attention Is All You Need', arXiv 1706.03762 v1, 12 June 2017, section 2 'Background', p. 2*

> As printed in the source: “To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using RNNs or convolution.”

To their knowledge, the first to rely *entirely* on it — not the first to have it.

Third, the part that mattered, and it is not the score. In a recurrent layer, relating two distant positions takes a chain of steps, one for each position in between. In a self-attention layer the connection is direct: the paper's sentence is that such a layer connects all positions with a constant number of sequentially executed operations. That is a claim about steps, not about speed, and not about using a distant word well.

The conference reviews are public, and one of them asked for something the paper never supplied.

> It would be good to see this empirically validated by evaluating performance on long sentences specifically.

> — *NeurIPS 2017, the published reviews of 'Attention is All you Need' (paper 3058), Reviewer 2, under 'Weaknesses'; https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Reviews.html*

Nine years late, a crude relative of that test: push the deciding word eight thousand words from the question, and its effect on the scores is about what it is at thirty-two words. The answer, though, stops being right long before that.

So why five more years before you heard of it? It did not sit idle. The BERT paper appeared in October twenty eighteen, using the architecture to read rather than write. A year later — three years before ChatGPT — Google said it was putting BERT into Search: into ranking, where it would help with about one search in ten in US English, and into featured snippets. The post described the family this way.

> models that process words in relation to all the other words in a sentence, rather than one-by-one in order.

> — *Pandu Nayak (Google), 'Understanding searches better than ever before', The Keyword, 25 October 2019, section 'Applying BERT models to Search'; https://blog.google/products-and-platforms/products/search/search-language-understanding-bert/, read 7 September 2026*

True of BERT, which can use both sides of a position. The model that writes your answer cannot, and that is the restriction from beat three.

People drew those numbers, then argued about the picture. In February twenty nineteen Sarthak Jain and Byron Wallace posted *Attention is not Explanation*, having found that very different weights could often be built that gave the same predictions. Their abstract ends like this.

> Our findings show that standard attention modules do not provide meaningful explanations and should not be treated as though they do.

> — *Sarthak Jain and Byron C. Wallace, 'Attention is not Explanation', arXiv 1902.10186 v1, 26 February 2019, abstract*

Six months later Sarah Wiegreffe and Yuval Pinter replied — *Attention is not not Explanation* — that such a claim depends on your definition of explanation. Then the fact that settles how much either tells you about the machine you use, checkable in fifteen seconds with find-in-page: across both papers the words "transformer" and "self-attention" appear zero times. Both are about attention inside recurrent models.

So I put the smallest version of that question I could to a model published this year, ranking every earlier word twice: by the attention the last position pays it, and by how much the answer moves when the word is taken out of what may be attended to. Those two orderings barely line up — a rank correlation of about zero point zero seven — and the words drawing most attention included a colon and the prompt's first token. That settles nothing in general. It is enough for this: a weight and an effect are different quantities.

Two things, each on a page, on a date.

One. In the open-weight models whose configuration files describe their layers, the all-pairs comparison is now the minority. I opened three on the seventh of September: Kimi K three designates twenty-four full-attention layers of ninety-three; Qwen three point eight, sixteen of sixty-four; G L M five point three Flash, none of forty-five. Those labels are the vendors' own, no independent audit turned up, and a count of layers is not a count of the work they do. And the comparison did not go away — in a sparse layer it still runs, as the step that *chooses* which words get attended to, which the IndexCache paper from March says keeps the very cost the sparse layer was built to escape.

Two. On the twenty-eighth of August, four researchers reported that a plain sliding window did as well or better than the linear-attention models they tested — models converted after training, not built that way. Watch whether anyone tries that on the second kind.

One idea, three moves. **A word becomes a position. Attention adds to that position using the other words. A layer does it again.**

And if you keep one image, keep the chain becoming a direct connection whose length no longer grows with the gap. The word attention names a weighting calculation; what made it matter was using that calculation everywhere, until distance stopped counting.

Then the honest part. Anthropic's interpretability team publishes on a site called the Transformer Circuits Thread. This is the first line on it.

> A surprising fact about modern large language models is that nobody really knows how they work internally.

> — *Anthropic, Transformer Circuits Thread, the opening paragraph of the site's front page; https://transformer-circuits.pub/, read 7 September 2026*

The arithmetic of a layer can be written out exactly. What a layer has learned to do with it is largely an open question.

Tomorrow: how those scores turn into one chosen word — and why the same question, asked twice, can come back different.

That was day four. Thank you for listening.

---

## Sources (15)

- Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin — *Attention Is All You Need* — 12 Jun 2017 (v1)
- NeurIPS 2017 — published peer reviews of that paper — Dec 2017
- Bahdanau, Cho, Bengio — *Neural Machine Translation by Jointly Learning to Align and Translate* — 1 Sep 2014 (v1)
- Jurafsky and Martin — *Speech and Language Processing*, 3rd edition draft — released 19 Aug 2026
- Jain and Wallace — *Attention is not Explanation* — 26 Feb 2019 (v1)
- Wiegreffe and Pinter — *Attention is not not Explanation* — 13 Aug 2019 (v1)
- Pandu Nayak, Google — *Understanding searches better than ever before* — 25 Oct 2019
- Devlin, Chang, Lee, Toutanova — *BERT* — 11 Oct 2018 (v1)
- Bai, Dong, Jiang, Lv, Du, Zeng, Tang, Li — *IndexCache* — 12 Mar 2026 (v1)
- DeepSeek-AI — *DeepSeek-V3.2* — 2 Dec 2025 (v1)
- Qiu and colleagues — *On the Design of Qwen3.8-Next Architecture* — 31 Aug 2026 (v1)
- Jolicoeur-Martineau, Sukthanker, Cameron, Gervais — *Sliding-window beats linear attention* — 28 Aug 2026 (v1)
- Vendor configuration files: `moonshotai/Kimi-K3`, `Qwen/Qwen3.8-27B`, `zai-org/GLM-5.3-Flash` — read 7 Sep 2026
- Anthropic — Transformer Circuits Thread — read 7 Sep 2026
- Direct measurement, one desktop machine — 7 Sep 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
