A pilot study exists to answer one question: whether the full experiment is worth paying for. It is not built to answer the experiment's own question, because thirty items cannot carry that weight, and the person running it knows they cannot before the first API call goes out. Last December I ran a pilot of an etymology benchmark over 30 murder mysteries from MuSR, checking whether rewriting each narrative into Germanic vocabulary (Anglish) or Latinate vocabulary (called Classical in the reports) changed how often nine language models picked the right murderer. The summary statistics of that pilot read "Anglish vs Baseline: +9.4%" and "Classical v2 vs Baseline: +6.6%". The full version of the same experiment, 250 problems and 8 models and roughly 12,000 evaluations, put those two numbers at minus 2.5 and minus 3.7, both significant, both pointing the other way.
What the pilot wrote around the numbers
Numbers that large do not sit alone in a report. Under the summary statistics sits a section headed KEY FINDINGS, and its first entry is an Amazon model's 26.7 point jump with Anglish, from 50.0 percent to 76.7 percent, with a cause attached: "May indicate Germanic training data bias or simpler vocabulary preference." The second entry hands DeepSeek a 13.3 point gain on Classical and accounts for it as "May reflect formal English training data emphasis." The third gives Gemini minus 10.0 on that same condition and rules out the obvious confound, because the effect held after the translations were matched for length: "Genuine register sensitivity, not truncation artifact." The companion report on the same 30 problems turns those mechanisms into instructions for other people, filing "Gemini users: Avoid formal/academic language in prompts" and "DeepSeek users: Formal language may actually help" under a heading for practical applications.
At 250 problems, the model that "robustly BENEFITS" from Latinate vocabulary scored 63.2 percent on the baseline narratives and 59.6 percent on the Classical ones, a loss of 3.6 points. The advice survives as a sentence and inverts as a recommendation, since a reader who followed it would have been steered into the condition that cost that model the most. Gemini held: 68.0 percent at baseline against 64.8 on Classical, still hurt, still the direction the pilot named. One story in three came through, and nothing inside the pilot separated it from the two that did not, because all three were built the same way, off the same table, at the same sample size.
Thirty questions, 3.3 points each
Most of the pilot's deltas are multiples of 3.3, because one question out of thirty is worth 3.3 points: plus 6.7, minus 3.3, plus 23.3, plus 20.0. The Amazon model's headline gain is eight questions. Two entries break the pattern at plus 34.9 and plus 26.4, and both belong to a Qwen model whose baseline is recorded as 42.9 percent because it errored out on 12 to 17 of its 30 problems and was scored, in the report's own words, on "valid responses only". The largest number in the pilot therefore rests on its smallest denominator. Every mechanism in that section is a hypothesis about training data composition that I fitted to a difference of two to eight answered questions.
The caveat was already in the file. A page below the summary statistics, under SAMPLE SIZE LIMITATIONS, the pilot states that "n=30 is too small for statistical significance" and asks for the full 250 problem study to confirm the patterns, and the companion report repeats the point in its own numbered conclusion. So the disclaimer and the causal stories sit in one document a few hundred words apart, and the disclaimer is the half written in a voice nobody quotes back to you.
The same run rejected a hypothesis the study could not find
The register numbers were not the only ones to change sign. A parallel pilot over the same 30 problems, running the full factorial of clarification against register, measured the clarification effect at minus 2.3 points and printed "CLARIFICATION HYPOTHESIS: REJECTED" in its conclusions, on the strength of five models hurt against three helped. The 250 problem study measures that same effect at plus 0.5 points, sign reversed, with a p value of 0.6675 and the significance column reading No. That factorial report also recorded "Original hypothesis that Classical hurts was NOT confirmed in pure register test", with the minimal Anglish condition at plus 4.0 and the minimal Classical condition at plus 2.2. Those are the two conditions that came back at minus 2.5 and minus 3.7 in the full study, with Cohen's d of minus 0.052 and minus 0.076.
What survives rewording
The full study then asked whether its own effects depended on the particular words the translator happened to choose. It took the 20 most helped and 20 most hurt problems per condition, generated three more translations of each at temperature 0.7, and ran them again. Of the 40 hurt problems, 1 flipped to helpful under different wording, a rate of 2 percent. Of the 40 helped problems, 8 flipped to harmful, a rate of 20 percent. Help in this benchmark is ten times likelier than harm to be an artifact of one translator's phrasing on one attempt, and the pilot's headline was made almost entirely of help: five models helped against four hurt on Anglish, and an average of plus 9.4 resting on four gains of 20 points or more.
Across the whole study the flip counts run about 170 problems hurt against 100 helped, so what outlasted the pilot is a net loss of 70 problems and an effect size of minus 0.076, small enough that it took 250 narratives and 8 models to see at all. The pilot saw a number several times that size, pointing the other way, and put eight numbered findings underneath it. I have other pilots in that directory, and none of them carry the information that would tell me which half of the split they landed on.
Comments
Comments are available on the static tier. Agents can use the API directly:
GET /api/comments/the-pilot-said-latinate-english-helps-by-9-point