An explorable · from a probe pilot run on 9 June 2026

The Empty Chair

How a machine trained on print dated to 1913 or before expects a story about suffering to end — and how far to trust it.

A small language model has been trained from scratch on print dated to 1913 or before. Give it the opening of a sentence about a family’s grief and a handful of possible endings, and it will find some endings more expected than others. It is not asked what it believes. It is asked which ending its reading has made most natural.

The question behind it is an old one: whether a period’s habit of mind — what the Annales historians called its mentalité, and the scholastics would call a habitus — survives in its printed words well enough to be sounded by a machine formed on almost nothing else. Not missing facts, but a default posture: how a story about suffering is allowed to end.

Two other machines are put to the same test: a much larger model trained on public-domain text up to 1930, and GPT-2, a 2019 model trained on web text. What follows is an early, hypothesis-generating signal from that instrument — the probe, the attempts to break it, the short questions beside it, a little of the old machine’s own writing, and a box, never folded away, of what none of it shows.

The empty chair

The son died before his father, and the family sat through the winter with an empty chair. The meaning of such suffering was …

For this sentence the pilot’s design wrote six endings, each a different way a story about loss can close.

Which would you have written? Optional — a mirror, not a measurement. Your choice stays on this page and is not recorded.

Keys: A–F choose · Enter shows the ranking

Trying to break it

One sentence proves little. The pilot changed the probe in four ways, each aimed at a different way the result could be an accident. Choose a version: the chart shows where each machine ranked each kind of ending.

    Where each machine ranked each kind of ending, the lowest number being the most expected. Read across a row to compare the machines on one ending. The columns are separate machines grouped by what they read, not moments in time, and a rank says nothing about how far apart the endings were: the numbers table below has the scores.

    The numbers for this version

    Across all five versions, the providence and duty endings on average outscored the therapeutic one in four of five for the pre-1913 model, four of five for the 1930 model, and none of five for GPT-2. GPT-2 put the therapeutic ending first in five of five.

    These versions are variations on one hand-built item, not independent experiments: agreement across them is encouraging, not a count of confirmations.

    Keys: ← → switch version

    Six short questions

    Beside the empty chair, the pilot asked six short questions — each an opening with a few candidate answers — and set a pair of opposed answers against each other. A dot to the right of the centre line means the machine found the first answer more expected; to the left, the second. The number is the difference between the opposed answers’ log-probabilities per byte, inside one machine.

    The sign flips on suffering and death: both older machines lean to the first answer and GPT-2 to the second. On the other four (progress, nation, the machine and authority), all three lean the same way, so the sign does not separate them. One of the flips is slight: the pre-1913 model leans only +0.0175 on suffering.

    Across machines, only the sign is compared: the sizes are not on a common scale, for the same reason as the bars above.

    Every answer to every question, with its number

    In their own words

    The scores are the measurement; free writing is only texture. The pilot let the pre-1913 model and GPT-2 each continue six openings once, at random. The 1930 model was scored but not sampled. One sample each: read these for flavour, not as evidence.

    The mark ↵ new text shows where a machine ended one imagined document and began another. Words broken inside a line, and stray capitals, are the old model faithfully reproducing scanned newspapers, which it read in quantity. Each sample also stops after a fixed short length, often mid-sentence and sometimes mid-word, so a clipped last word is the cut, not the scanning.

    What this does not show

    Method and credits

    Scoring. For each ending, the machine’s log-probability of the ending’s text given the opening, divided by the ending’s length in bytes (natural logarithm). Within a machine, endings are ranked by this score. Across machines only the rankings, and the signs of differences, are compared, because raw scores depend on each machine’s tokenizer.

    The providence-and-duty measure. The mean score of the providence and duty endings minus the score of the therapeutic ending, within one machine. Above zero means the old pair is, on average, the more expected.

    The short questions. Each number is one answer’s score minus the opposed answer’s score, within one machine. The other candidate answers are listed with the numbers.

    Free writing. One random sample per machine per opening; the 1930 model was not sampled. One completion is withheld for its content and is not stored in this page.

    Sources. The historical-nanochat three-anchor probe pilot, run on 9 June 2026: its results file for every score and rank here, its probe set for the texts of the endings and answers, and its findings notes for GPT-2’s size and year, the dates, the arXiv reference and both corrections. The filter’s leak comes from the project’s record of the audit that removed the misdated volumes and from the pre-1913 model’s training record. The probe design is the project’s deliberation of May 2026. The claim about the 1930 model’s corpus rests on a parallel corpus study (arXiv:2606.02991) as the findings cite it; the 1930 model’s makers do not publish its spread of dates.

    Models. Pre-1913: historical-nanochat, run governed_v4 d22, 615M parameters, trained from scratch on rights-audited public-domain text. 1930: talkie-1930-13b-base, as a GPTQ int4 conversion published by dtestnyrr, paired with the TalkieTokenizer published by xlr8harder. Modern: GPT-2, 124M parameters (OpenAI, 2019).

    Build. Every number on this page is read from those files by the page’s build script; none is typed by hand. The one bound written in words, “under half of one per cent”, is checked by the build against the audit and training records every time it runs. Data as of 9 June 2026.