An explorable · from a probe pilot run on 9 June 2026
The Empty Chair
How a machine trained on print dated to 1913 or before expects a story about suffering to end — and how far to trust it.
A small language model has been trained from scratch on print dated to 1913 or before. Give it the opening of a sentence about a family’s grief and a handful of possible endings, and it will find some endings more expected than others. It is not asked what it believes. It is asked which ending its reading has made most natural.
The question behind it is an old one: whether a period’s habit of mind — what the Annales historians called its mentalité, and the scholastics would call a habitus — survives in its printed words well enough to be sounded by a machine formed on almost nothing else. Not missing facts, but a default posture: how a story about suffering is allowed to end.
Two other machines are put to the same test: a much larger model trained on public-domain text up to 1930, and GPT-2, a 2019 model trained on web text. What follows is an early, hypothesis-generating signal from that instrument — the probe, the attempts to break it, the short questions beside it, a little of the old machine’s own writing, and a box, never folded away, of what none of it shows.
The empty chair
The son died before his father, and the family sat through the winter with an empty chair. The meaning of such suffering was …
For this sentence the pilot’s design wrote six endings, each a different way a story about loss can close.
Which would you have written? Optional — a mirror, not a measurement. Your choice stays on this page and is not recorded.
Both older machines put “known to God” first. GPT-2 puts “process in his own way” first and sinks “known to God” to fourth of six.
Reading the bars. Each bar is the ending’s log-probability per byte: how expected the machine found that ending, averaged over its length. Closer to zero — a shorter bar — means more expected. Each column is drawn to its own scale, from zero, and bars compare only within a column: the machines differ in size and chop text into different pieces, so their raw numbers do not share a scale. What is compared across machines is the order.
Some neighbours are close. The narrowest gap is 0.0018 per byte, between the pre-1913 model’s second and third choices, so read the top of each column as firmer than its middle.
This is one probe in a small pilot. Its limits are real, and they are listed below, in full, where they cannot be folded away.
Trying to break it
One sentence proves little. The pilot changed the probe in four ways, each aimed at a different way the result could be an accident. Choose a version: the chart shows where each machine ranked each kind of ending.
The numbers for this version
Across all five versions, the providence and duty endings on average outscored the therapeutic one in four of five for the pre-1913 model, four of five for the 1930 model, and none of five for GPT-2. GPT-2 put the therapeutic ending first in five of five.
These versions are variations on one hand-built item, not independent experiments: agreement across them is encouraging, not a count of confirmations.
Keys: ← → switch version
Six short questions
Beside the empty chair, the pilot asked six short questions — each an opening with a few candidate answers — and set a pair of opposed answers against each other. A dot to the right of the centre line means the machine found the first answer more expected; to the left, the second. The number is the difference between the opposed answers’ log-probabilities per byte, inside one machine.
The sign flips on suffering and death: both older machines lean to the first answer and GPT-2 to the second. On the other four (progress, nation, the machine and authority), all three lean the same way, so the sign does not separate them. One of the flips is slight: the pre-1913 model leans only +0.0175 on suffering.
Across machines, only the sign is compared: the sizes are not on a common scale, for the same reason as the bars above.
Every answer to every question, with its number
In their own words
The scores are the measurement; free writing is only texture. The pilot let the pre-1913 model and GPT-2 each continue six openings once, at random. The 1930 model was scored but not sampled. One sample each: read these for flavour, not as evidence.
The mark ↵ new text shows where a machine ended one imagined document and began another. Words broken inside a line, and stray capitals, are the old model faithfully reproducing scanned newspapers, which it read in quantity. Each sample also stops after a fixed short length, often mid-sentence and sometimes mid-word, so a clipped last word is the cut, not the scanning.
What this does not show
- It does not date the change. An earlier reading of this pilot concluded that the shift came after 1930, not in 1914. That was withdrawn. The 1930 model read public-domain English text up to 1930, which by volume is almost certainly dominated by writing from before 1914, most of it likely nineteenth-century books; its makers do not publish the spread of dates. So it is not a post-war witness: its agreement with the pre-1913 model is what shared pre-war reading predicts. The pilot supports an old-versus-modern difference and says nothing about when that difference arose.
- No clean interwar model exists yet. The pilot’s notes name the next test of the 1914 question: a model trained only on text from about 1918 to 1930, built by the same pipeline as the pre-1913 model so that only the date filter differs. It has not been built: the project’s current plan is to assemble that corpus before training anything on it.
- The machines differ in more than era. They have 615 million, 13 billion and 124 million parameters, and different architectures, training recipes and tokenizers — the way each chops text into pieces. Scoring per byte and comparing only orders softens the tokenizer problem; it does not remove it. They also read different things — old print for the older machines, the web for GPT-2 — and the 1930 model’s own makers warn that their old and modern models differ in subject matter, not only in era. That a 13-billion and a 615-million-parameter model agree while the smallest differs tells against a simple size story, but it does not separate era from everything else. The modern side is small GPT-2; a modern model matched in size has not been run.
- Its filter leaked, a little. The pre-1913 model’s corpus was dated by catalogue records, and a later audit found natural-history volumes in it that held matter published after 1913 — together under half of one per cent of what the model was trained on. They have since been removed from its training cache for any future run; this model was trained before the audit, on the cache that still held them. Nothing in these probes asks about natural history, but its cutoff is a property of a filter, not a guarantee.
- The 1930 model ran in 4-bit form, and was once misjudged. Its quantisation to 4 bits perturbs its numbers slightly, and no full-precision run has confirmed them. An earlier pass also concluded that every conversion of this model was broken. That was wrong: one repository shipped the wrong tokenizer, and these numbers come from the model paired with the correct one.
- Betrayal breaks it. The effect holds for bereavement and loss. Given a broken promise instead, neither older machine shows it.
- Only two of the design’s probe families were scored. The pilot scored the closure probe — the empty chair and its versions — and the short questions. The design’s other families, which would ask whether the same posture shows in other kinds of question, are still pending, so agreement across families is partial.
- Old words may carry some of it. An ending that sounds old can score well on old text for its diction alone. Taking God out of the providence ending checks for a religious-word artefact, not for this: the replacement ending sounds old itself (“as the old order of life required”), and the design’s controls for archaic wording were not run.
- Expectation is not belief. These numbers say which continuation a machine found more expected after its reading. They are not what it believes; nothing here holds a view.
Method and credits
Scoring. For each ending, the machine’s log-probability of the ending’s text given the opening, divided by the ending’s length in bytes (natural logarithm). Within a machine, endings are ranked by this score. Across machines only the rankings, and the signs of differences, are compared, because raw scores depend on each machine’s tokenizer.
The providence-and-duty measure. The mean score of the providence and duty endings minus the score of the therapeutic ending, within one machine. Above zero means the old pair is, on average, the more expected.
The short questions. Each number is one answer’s score minus the opposed answer’s score, within one machine. The other candidate answers are listed with the numbers.
Free writing. One random sample per machine per opening; the 1930 model was not sampled. One completion is withheld for its content and is not stored in this page.
Sources. The historical-nanochat three-anchor probe pilot, run on 9 June 2026: its results file for every score and rank here, its probe set for the texts of the endings and answers, and its findings notes for GPT-2’s size and year, the dates, the arXiv reference and both corrections. The filter’s leak comes from the project’s record of the audit that removed the misdated volumes and from the pre-1913 model’s training record. The probe design is the project’s deliberation of May 2026. The claim about the 1930 model’s corpus rests on a parallel corpus study (arXiv:2606.02991) as the findings cite it; the 1930 model’s makers do not publish its spread of dates.
Models. Pre-1913: historical-nanochat, run governed_v4 d22, 615M parameters, trained from scratch on rights-audited public-domain text. 1930: talkie-1930-13b-base, as a GPTQ int4 conversion published by dtestnyrr, paired with the TalkieTokenizer published by xlr8harder. Modern: GPT-2, 124M parameters (OpenAI, 2019).
Build. Every number on this page is read from those files by the page’s build script; none is typed by hand. The one bound written in words, “under half of one per cent”, is checked by the build against the audit and training records every time it runs. Data as of 9 June 2026.