2026-08-29|10 min read
375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
Method and raw data: Two Viral Prompt Wrappers: Method, Arms and the Full Grid
2026-08-15|3 min read
104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.
Method and raw data: Four Models on a Retrieval Contract: Method, Both Realizations, and What the Instrument Cannot Say
2026-08-13|5 min read
A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7.
2026-08-13|5 min read
An experiment whose entire premise is a controlled information boundary shipped a finished manuscript and left the one instrument that would measure the boundary with no records in it.
2026-08-11|2 min read
The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.
Method and raw data: Model Identity Verification: Methodology and Raw Data
2026-08-11|3 min read
Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.
Method and raw data: Harness Replication: Method, Thirteen Cells and the Artefacts Caught
2026-08-11|3 min read
A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
Method and raw data: The Reasoning Floor: Methodology, Cell Data and Control Probe
2026-08-11|3 min read
Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant.
Method and raw data: Prompt Stack Trim: Method, Taxonomy and the Ranked Proposal
2026-08-11|3 min read
Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.
Method and raw data: Speech to Text Bake-off: Methodology, Reference Design and Full Tables
2026-08-11|3 min read
Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.
Method and raw data: DeepSeek V4 Flash against GPT-5.6 Luna: Methodology, Results and Retraction
2026-08-11|3 min read
GPT-5.6 Luna and GPT-5.3 Spark at effort low, given a plain-language prompt, against hand-written keyword and regex classifiers on four real classification sites, independently adjudicated, with a split verdict and a batching correction.
Method and raw data: Semantic against Keyword: Method, Adjudication and Per-site Verdicts
2026-08-06|3 min read
A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted.
Method and raw data: Cache-Read Coefficient: Methodology and Raw Data