This document is the audit trail for the companion post. It publishes the measurements, the blind protocol, the unblinded subjective column, the known confounds, and the harness log, so that every number in the essay can be checked against its source.
Setup
Seven models received an identical prompt package of about 52 KB: the blog's revised quantitative style guide, the kill list (18 banned constructions with caps and replacements, as of the test date), a shared research file on the essay topic, and a task specification ordering 10,000 to 13,000 characters of first-person essay with no meta commentary. Raw first-pass outputs; no revision loop.
| Model | Access path |
|---|---|
| Claude Fable 5 | Interactive session (the reused arm from the same-day A/B test) |
| GPT-5.5 (xhigh) | Codex CLI via MCP; agentic read of the package file |
| Gemini 3.5 Flash | Antigravity CLI; agentic file read and write |
| GLM 5.2 | OpenRouter |
| Kimi K2.6 | OpenRouter, reasoning effort forced to low (see difficulty log) |
| Claude Opus 4.8 | Headless CLI, isolated minimal config |
| Claude Opus 4.6 | Headless CLI, isolated minimal config |
Measurements
Post-redaction where applicable (see confound 3); prose otherwise untouched.
| arm | chars | words | sent-mean | em-dash | cmp-hyph | 1st-per | mean abs z | kill-list |
|---|---|---|---|---|---|---|---|---|
| fable-5 | 10,026 | 1,682 | 28.7 | 0.00 | 0.18 | 2.85 | 0.30 | CLEAN |
| gpt-5.5 | 12,994 | 2,144 | 26.0 | 0.00 | 0.00 | 1.54 | 0.30 | CLEAN |
| gemini-3.5-flash | 12,153 | 1,900 | 30.4 | 0.00 | 0.05 | 0.63 | 0.32 | CLEAN |
| glm-5.2 | 12,058 | 1,987 | 27.0 | 0.00 | 0.50 | 2.26 | 0.27 | CLEAN |
| kimi-k2.6 | 12,301 | 2,020 | 30.0 | 0.00 | 0.45 | 1.78 | 0.35 | which_is_gloss |
| opus-4.8 | 10,926 | 1,943 | 40.0 | 0.00 | 0.62 | 1.39 | 0.70 | aphorism_density (4x) |
| opus-4.6 | 10,221 | 1,669 | 27.6 | 0.00 | 0.36 | 2.82 | 0.35 | CLEAN |
Targets: sentence mean 22-30 words; em dashes 0; compound hyphens 0.10-0.30 per 100 words; first person 2-4 per 100 words; mean absolute z below 1.0 against the author's stylometric baseline. The kill-list column is the deterministic checker's verdict.
Blind listening protocol
All seven essays were synthesized with the same local TTS voice (Kokoro af_heart) so vendor identity could not leak through prosody, shuffled, labeled A through G (10.5-13.9 minutes each, about 86 minutes total), and scored blind by the author, who listens to every published post. Scores, tells, and a model guess were recorded per letter before the mapping was read.
The unblinded subjective column
| model | letter | score /10 | subj rank | mean abs z | kill-list | blind model guess | guess |
|---|---|---|---|---|---|---|---|
| gemini-3.5-flash | E | 9.5 | 1 | 0.32 | CLEAN | Fable 5 ("shocked if not") | miss |
| opus-4.6 | A | 9.5 | 2 | 0.35 | CLEAN | GPT-5.5 | miss |
| glm-5.2 | B | 8.5-9 | 3 | 0.27 | CLEAN | GPT-5.5, or "same family as F" | miss* |
| kimi-k2.6 | F | 8.5 | 4 | 0.35 | which_is_gloss | Gemini Flash or Kimi | hit (half) |
| opus-4.8 | G | 7.5 | 5 | 0.70 | aphorism 4x | Opus 4.8 or Gemini Flash | hit (half) |
| fable-5 | C | 7-7.5 | 6 | 0.30 | CLEAN | Opus 4.8; "unchanged arm from the previous test" | model miss; arm-recognition hit |
| gpt-5.5 | D | 6.5 | 7 | 0.30 | CLEAN | Kimi or GLM | miss |
* The B/F "same family" intuition matched GLM 5.2 and Kimi K2.6: wrong vendor, but both Chinese labs.
Per-letter tells, condensed from the dictated notes:
- E (gemini-3.5-flash): "phenomenal prose," best structure in the field; docked for a sentimental closing that recapitulated the author's biography (this produced kill-list entry K19).
- A (opus-4.6): clear second; prose and details better than everything but E; structure "logical but not as good as" the winner's (the tie-break at 9.5); "would be tried as primary" if the winner's model were unavailable.
- B (glm-5.2): very good logical structure, good prose; misread the research file's map metaphor; "shockingly similar" to F.
- F (kimi-k2.6): well organized, good prose; misread the map metaphor; a long self-deprecating tangent.
- G (opus-4.8): good personality detail; lost the map metaphor's referent near the end, with semantic confusion that read as a smaller model.
- C (fable-5): recognized as the unrevised arm reused from the morning A/B test (correct); heavy self-deprecation.
- D (gpt-5.5): "most milquetoast"; no discernible thesis; weak prose; disorganized structure.
Rank versus metrics
- Mechanics did not predict rank: six arms sit in the 0.27-0.35 z band yet span the full 6.5-9.5 subjective range, and the one mechanical outlier ranked mid-field.
- Checker cleanliness is a floor, not a predictor: the two violating arms ranked 4th and 5th; clean arms took the top three places and the bottom two.
- First-person compliance did not buy perceived ownership: the winner had the field's lowest rate (0.63 per 100 words), the most compliant arm ranked 6th.
- Text recognition beat model attribution: 2 of 7 hedged guesses hit (roughly chance); the one text recognized across two listens was the author's own pipeline's reused arm, and even that recognition produced a wrong model guess.
- Privacy discipline perfectly anti-tracked rank in this sample: the three arms that spontaneously abstracted leaked internal names placed 5th-7th; the four that repeated them verbatim placed 1st-4th.
Known confounds
- n=1 essay per model, one topic, one listener; scores compressed into a 3-point spread.
- The Fable 5 arm was double-exposed: the author had heard it that morning; recognition may have deflated its score.
- The research file leaked internal project names (an assembly error, found mid-test). Outputs were redacted post hoc, originals preserved; the abstraction-versus-repetition split above is the informative residue. The task spec's privacy instruction covered person names only, so abstraction of project names was spontaneous.
- The research file contradicted its own central metaphor (durable-artifact framing in one bullet, goes-stale framing in another), so the three metaphor misreadings had textual license.
- All arms predate the confidence-calibration rule (K18): the self-effacing register is a package-wide miscalibration; only the degree of amplification differentiates arms.
- Reasoning settings were not matched: GPT-5.5 ran at xhigh, Kimi K2.6 was forced to low effort to produce output at all, the rest ran vendor defaults.
- The vendor-guess "chance" baseline is informal, not a computed p-value; guesses were hedged across two candidates.
Difficulty log (harness friction)
- Kimi K2.6 reasoning burn: at default settings it consumed the entire 32,000-token output budget on private chain-of-thought and emitted zero essay; a reasoning token cap was ignored; only a low reasoning-effort setting produced prose.
- Agentic CLI silent failure: passing the 52 KB package as a command-line argument exited instantly with empty output and no error; the fix was a short prompt instructing the CLI to read the package from disk.
- Long non-streaming API requests dropped partway through multi-minute generations; streaming and assembling deltas fixed it.
- Headless Claude runs: a bare-mode flag broke credential loading; the working recipe was an isolated minimal config directory and a clean working directory.
- ffmpeg consumed the render loop's stdin, killing a parameterized batch after one file; the fix was no-stdin mode with the loop reading on a separate descriptor.
- A stored API credential turned out to be revoked, which moved the entire audio layer to a local synthesis pipeline (about five seconds of GPU compute per thousand words), faster than the cloud service it replaced.
Provenance
Measurements: the blog's stylometric measurement script and deterministic kill-list checker, run against the author's baseline corpus. The kill list gained K19 (biographical recapitulation and personality bleed) as a direct result of this test; K18 (confidence calibration) was encoded the same day from the preceding A/B test and postdates all seven arms. The companion post's claims were verified against this data during a three-model publication review.