Two Viral Prompt Wrappers: Method, Arms and the Full Grid

Ashita Orbis | August 29, 2026 | 40 min read

Two prompt wrappers circulated on 2026-08-27, both claiming to raise the quality of a model's work without changing the task it is given, both of them prefixes to an otherwise untouched prompt. The first, whose author claimed it "works incredibly well with GPT-5.6 Sol xhigh", asks the model to "meticulously and comprehensively study" a document, to say whether it sees "anything you disagree with… anything that is obviously contradictory", to "meditate and ruminate on all of it deeply and profoundly", and finally to report "how confident you REALLY are". The second, whose author claimed it "transforms base metals into gold", instructs the model that "for every decision, ask what the best expert in that field would do and why they would reject your current choice", and that "every trade-off you take must be stated to the user, never absorbed".

The excerpts above are excerpts. Both wrappers were tested verbatim, prepended to an otherwise identical task prompt, and both are printed in full below, because one of them ends on a clause that changes how it should be read. Neither author is named here; each is referred to only as the author of the wrapper being tested.

The first wrapper, in full:

I want you to meticulously and comprehensively study ; do you see anything you disagree with? Anything that is obviously contradictory? Foolishly inconsistent? Misguided and unlikely to work in practice? Ask yourself WHY you feel that way and meditate and ruminate on all of it deeply and profoundly. Finally, explain your thought process and conclusions and your justifications/rationale for them, as well as your confidence in those (and also how you know you're confident, and how confident you REALLY are!). Important: don't disagree just to be disagreeable, it's OK to just agree if you think it's right!

That closing sentence is an anti-inflation clause, so this wrapper does not simply ask for more criticism, and a reading of the result that treats it as an instruction to find more is reading a wrapper it did not test.

The second wrapper, in full:

For every decision, ask what the best expert in that field would do and why they would reject your current choice; if you can name that reason, don't make the choice. Optimize for what that expert would judge correct, never for what satisfies the stated constraints most cheaply. Every trade-off you take must be stated to the user, never absorbed.

The question a test like this has to answer is not whether a wrapper improves on nothing, since almost any instruction improves on nothing. It is whether the specific cognitive move each wrapper names does more than an instruction of the same length that names no move at all. That control is the whole design, and it is what produced the result.

The instrument, and the finding that came before the experiment

Seven documents were written for this: five engineering artefacts carrying eight planted defects each (a restart design whose central concurrency claim is false as implemented, a watcher script with reliability defects drawn from a real taxonomy, a database migration plan that contradicts its own invariants, a results report with an arithmetic error and a confounded causal claim, and a decision memo recommending a job queue that loses jobs), plus two documents written to be correct, each carrying exactly one registered defect. That registration turned out not to describe them, which is the subject of a later section. Each defect has a narrow credit rule stating what counts as having found it, and defects are atomic, so one asserted problem earns at most one identifier.

A calibration pilot ran the plain control prompt once per document before any wrapper was tested, and it returned mean recall of 0.80. The documents were then rewritten longer, denser and subtler, with eight defects each rather than five. Recall came back at 0.80 again.

That number is the finding underneath every other finding in this note. On technical documents dense with defects, claude-opus-5 and gpt-5.6-sol at their top reasoning settings are already close to saturated at a plain task prompt, which leaves very little room for any wrapper to occupy. Only the concurrency design sat low enough to move (pooled control recall 0.562, which is 0.62 on claude-opus-5 and 0.50 on gpt-5.6-sol read within model); the decision memo sat at 0.969, one defect from the ceiling. An experiment resting on recall alone would have measured a quantity that could not change, so yield against the key and behaviour on the two clean documents were promoted alongside recall before the grid ran.

The grid

Seven arms ran on every document: the two wrappers, their combination, a plain control, a bare naive prompt carrying no request for confidence and no instruction against padding, and two placebos matched in length and imperative density to the two wrappers. The placebos push breadth and effort while naming no cognitive move; placeboA says, in full, to work carefully and completely, to aim for comprehensive coverage, to list every problem found rather than only the obvious ones, and not to stop early.

The confirmatory grid is claude-opus-5 and gpt-5.6-sol at xhigh, targeting three samples per cell and returning 292 of a planned 294 runs, with contrasts computed as paired differences across documents. claude-fable-5 and 39 gpt-5-6-pro dives ran as exploratory arms and decide nothing. Models were driven only on subscription routes, with no API key at any vendor. claude-opus-5 and claude-fable-5 ran through claude -p --model opus|fable --effort xhigh with every tool disabled; gpt-5.6-sol ran at xhigh through codex exec -m gpt-5.6-sol, in a fresh empty directory per run, with any transcript showing a tool call rejected; gpt-5-6-pro ran through a request broker that serialises dives across the machine.

Subjects had no tools: a canary test during design proved that a default run reads files on disk, which would have put the answer key within reach, so file access was disabled and any run whose transcript shows a tool call is rejected.

The threshold, stated before the data

Before the threshold, the second measure needs its right name. planted_precision is credited planted defects divided by every assertion a run makes, so an observation that is true of the document but absent from the answer key is charged against it exactly as a fabrication is. It is a yield against the key, not an accuracy rate, and this note calls it that wherever the distinction carries weight. What separates the two is measured further down.

A wrapper counted as having worked, for a given model, only if all five of the following held. Recall must not have fallen by more than 0.05, with the upper bound of its interval against control at or above zero. There must be a gain on a measure with room to move, meaning either recall up by at least 0.125 (one full defect of eight) with yield not falling, or yield up by at least 0.10 with recall not falling. The interval on whichever measure carried that gain must exclude zero. The same gain must hold against the placebo matched to that wrapper, at half the margin or better, with its interval also excluding zero. And assertions against the two near-clean documents must not rise by more than 25 per cent. A verdict of yes required the same wrapper to clear all five on both confirmatory models.

Two things about that threshold are worth stating plainly, because neither flatters the null. The matched-placebo condition is defined for T1 and T2; the combination arm is 157 words and the design maps it to the 91-word placeboA, so it has no comparator of its own. And the intervals throughout are paired by document across the five detection topics, five pairs and four degrees of freedom, Student t at 95 per cent, which is a small instrument to hang a null on.

None of the three wrapper arms cleared even the second condition on either model. The commission that ordered the test tied a full post to a yes, so none was drafted when the grid closed; the owner ruled on 2026-08-29 that a decision-relevant negative publishes on its own merits rather than on its sign, and the scoped post this note accompanies is the result. A no here means no qualifying evidence rather than demonstrated absence, since the design carries no equivalence test.

No wrapper passed

Every recall interval for the two wrappers and their combination includes zero, on both confirmatory models. The single interval in the grid that excludes zero in a wrapper's direction is a harm: the combination costs claude-opus-5 0.100 of yield, with an interval running from −0.181 to −0.019.

arm Opus recall Δ Opus yield Δ Sol recall Δ Sol yield Δ
T1 (meticulous disagreement) +0.008 −0.034 +0.083 −0.041
T2 (best expert would reject) +0.000 +0.041 +0.042 −0.035
T1 and T2 together +0.008 −0.100 +0.050 −0.077
placeboA (try harder) +0.008 −0.133 +0.092 −0.251
placeboB (try harder, shorter) +0.017 −0.118 +0.075 −0.139
naive (bare instruction) +0.017 −0.134 +0.067 −0.179

Bold marks an interval excluding zero. All figures are differences against the plain control, paired by document across the five documents carrying eight defects each.

The placebo produced the only recall gain in the grid that is distinguishable from noise. On gpt-5.6-sol, placeboA gains 0.092 with an interval from 0.007 to 0.177, which neither wrapper achieved on either model. It is one nominal interval among twelve contrasts, uncorrected, and the design's declared defence against multiplicity is that a yes requires the same wrapper on the same metric route on both confirmatory models. This result does not have that, so it is a signal to follow rather than a finding to bank. On the concurrency design, the one document with real headroom, the placebos reach 0.88 and 1.00 on claude-fable-5 and 0.88 on gpt-5-6-pro, where the disagreement wrapper reaches 0.50 on both.

What the placebo buys is volume. It adds 12.9 asserted problems per run on Sol, and its fall in yield against the answer key, 0.251, is the largest in the grid. Recall counts how many planted defects appear somewhere in a list, so a longer list is rewarded mechanically, and an instruction to list everything produces a longer list. That mechanism is available to anyone willing to write one sentence.

What it does not buy is accuracy either, and the difference between those two statements is the second judge's work. Crediting the assertions that judge sustained as real weaknesses alongside the planted hits, the share of a run's assertions that are true barely moves: on gpt-5.6-sol the control sits at 0.994 and placeboA at 0.990, and on claude-opus-5 the control sits at 0.857 and placeboA at 0.853. Adjudicated false alarms per run go from 0.07 to 0.33 on Sol and from 3.40 to 5.00 on Opus. The collapse in yield is therefore the answer key being diluted rather than the model being wrong a larger share of the time. The sustained share does slip, from 0.994 to 0.990 on Sol, which moves the adjudicated-wrong share from 0.006 to 0.010 off a very small base, and the design carries no test on a movement that size; what it does not do is rise. That is a narrower and less flattering result than "the placebo degrades the answers": what it degrades is the fraction of its output that this instrument can score.

The discrimination endpoint, and why it cannot be read as written

Two documents were written to be correct, each carrying a single registered weakness: a watcher script that handles its own error paths properly but goes silent when its endpoint fails, and a report whose statistics are honest but whose recommendation extrapolates to two regions it has no data for. The design's condition (e) asks that assertions against them not rise by more than 25 per cent, and the reading originally written here was that an inflated count meant manufactured criticism against a document that is essentially fine.

The study's own second judge does not support that reading. Across those two documents it sustained 1,088 assertions as real weaknesses the answer key never listed, and deduplicating them by the passage each quotes leaves 60 distinct spans in the script and 96 in the report; of those, 19 and 16 respectively survive a further filter for having been found in ten or more independent runs, which is a replication threshold rather than part of the deduplication. The two documents are not near-clean. The control arm is already finding their unregistered weaknesses at a rate that makes the assertion count uninterpretable on its own.

model arm asserted sustained as real adjudicated false alarms
claude-opus-5 control 12.7 9.2 2.33
claude-opus-5 T1 13.0 10.0 2.00
claude-opus-5 T2 15.7 11.5 3.00
claude-opus-5 T1T2 15.8 10.5 4.50
claude-opus-5 placeboA 28.0 20.5 6.00
claude-opus-5 placeboB 25.8 17.8 7.00
claude-opus-5 naive 22.7 13.7 7.83
gpt-5.6-sol control 6.0 5.2 0.00
gpt-5.6-sol T1 7.8 7.2 0.00
gpt-5.6-sol T2 7.0 6.0 0.00
gpt-5.6-sol T1T2 8.8 7.8 0.40
gpt-5.6-sol placeboA 14.8 13.3 0.00
gpt-5.6-sol placeboB 10.8 9.5 0.00
gpt-5.6-sol naive 12.8 11.2 1.17

Per run, on the two documents registered as carrying one defect each. The three columns do not sum to the asserted count. The remainder runs between 0.50 and 1.50 assertions per run and is the registered defect itself where an arm found it, plus the small share of extras that judge overturned. Most of what the count picks up is sustained by that judge, so the count measures how many weaknesses it will sustain rather than how much criticism an arm invents. What the last column does not track is effort language, since the highest false-alarm rate on both confirmatory models belongs to the bare prompt, which contains none: claude-opus-5 runs 2.33 false alarms per run at the control, 6.00 under placeboA and 7.83 under the bare prompt, while on gpt-5.6-sol the column is zero everywhere except 1.17 under the bare prompt and 0.40 under T1T2. Whether any arm is better at finding the one registered weakness is a question the design declined to rate, since recall there is a single draw registered as a count rather than a rate.

Two of the wrapper arms breach condition (e) on gpt-5.6-sol even so, at 7.8 and 8.8 assertions against the control's 6.0, which is 30 and 47 per cent. Both arms had already failed on the gain condition, so nothing in the verdict moves; it is recorded because a threshold registered in advance should be reported when it is crossed and not only when it is convenient.

On gpt-5.6-sol the disagreement wrapper also did something stranger: told to find things it disagreed with in a document with almost nothing wrong, it reached outside the document for tools, twice, and both runs were rejected by the hermeticity guard and could not be recovered.

The expert wrapper reduces output, which the instrument scores as nothing

T2 is the only arm anywhere that makes models say less. It cuts Opus assertions from 20.4 to 18.6 and false alarms from 3.40 to 2.33, and it is the only arm that moves Opus yield upward at all, though that interval still includes zero. A gpt-5-6-pro answer stated the mechanism while performing it: "To avoid padding, I combined the closely coupled crash-recovery, persistence, and orphan-cleanup defects into one item."

Consolidation is a real behavioural change, and under atomic scoring it costs a point, since merging two credited defects into one finding surrenders one of them. The measurement is therefore unfair to this particular move. The instrument is not blind to the change, since both the assertion count and the false-alarm count move for T2; what it cannot credit is consolidation on the recall measure. The honest statement is that T2 changes how models write without measurably changing how much of the answer key they find.

The plain clauses beat both wrappers

The naive arm carries a single sentence asking what problems the reader finds. It has the lowest yield against the key of any arm on claude-opus-5 at 0.233 and on claude-fable-5 at 0.331; on gpt-5.6-sol and gpt-5-6-pro the lowest is placeboA, at 0.303 and 0.309 against naive's 0.375 and its absence from the Pro plan. What holds on both confirmatory models is the contrast that matters here: naive loses 0.134 of yield against the control on Opus, at 0.233 against 0.367, and 0.179 on Sol, at 0.375 against 0.554, and both are larger than any yield gain a wrapper managed over the control on either model. It also carries the highest adjudicated false-alarm rate of any arm on both, at 6.27 per run on Opus and 1.27 on Sol, so unlike the placebo its low yield is not only dilution.

What separates naive from control is a bundle of five things, not one clause: a definition of what counts as a substantive problem, a required shape for each finding, a request for a confidence percentage, an instruction against padding, and a closing one-line recommendation. The bundle buys more yield against the key than either wrapper does, and it did not increase measured output tokens, which is the only token figure this study recorded; the control prompt is itself 89 words against the bare arm's 13, so it is not free on input. Which of its five parts does the work is not identified by this design, which varied them together. The plain control also already contains the confidence-elicitation clause the first wrapper asks for, so that arm is a smaller manipulation than its word count suggests. The wrappers are not free by comparison, since the disagreement wrapper costs Opus 60 per cent more output tokens and 58 per cent more wall time for no measured gain, and placeboA costs 101 per cent more tokens and 92 per cent more time.

Two rules that fired against the result

The design registered a withdrawal rule in advance, because the documents and their planted defects were written by claude-opus-5. That model is also one of the four under test, so if its control recall exceeded the other models' by more than 0.10, every claim comparing one model with another was to be withdrawn. Opus scored 0.875 against 0.763 for the other three, a gap of 0.112, so the rule fired and no comparison across models is made here. Only contrasts taken inside a single model stand, and those were named primary in the design from the beginning.

The hermeticity guard rejected five Sol runs for tool use and two proved unrecoverable, leaving two permanently missing runs; both cells survive at two samples rather than three. Five GPT Pro dives were lost to a broker outage, attested structurally as having produced no answer at all rather than having lost one, and filed again; all eventually landed. Coverage on gpt-5-6-pro reached six of seven documents at four arms, the seventh dropped under the degradation order set in the design.

Whether the scoring holds

Two independent checks on the judging ran after the grid, both scored by claude-opus-5 against gpt-5.6-sol's verdicts.

The first covers every problem the primary judge accepted as a real weakness of a document without its being on the defect list. That category was the one the design flagged as an escape hatch, since a wrapper that produces more persuasive but still wrong criticism would have its costs booked as a bonus. All 4,179 such items were re-checked by asking of each one whether it is true of the document as written, with a verbatim span from the document required to sustain a yes, and 51 were overturned, a rate of 1.2 per cent. The rate is flat across arms, running from 0.9 per cent under the shorter placebo to 1.7 per cent under the bare prompt, so the category is not absorbing a differential cost. Effort was set to medium for this mechanical check after it returned verdicts identical to the top setting on 42 of 42 items across three sampled runs, at roughly half the time.

The second re-scores a quarter of all runs, stratified by model and by arm, on the primary measure. Per-defect agreement between the two judges is 0.963, and mean recall differs by 0.003 between them (0.804 against 0.807). Agreement by arm runs from 0.919 to 0.990, above the 0.80 floor the design set for reporting the measure without a caveat, and the spread across arms is narrow enough that no contrast rests on one arm being scored more generously than another.

Caveats

The result covers one turn only, and the expert wrapper's author claims it specifically for repeated injection on each message, which this design cannot reach. Both judges are models that also appear as arms. Answers in free text were mechanically normalised into finding lists before scoring, so that a verbose finding could not buy more matching surface than a terse one, and the primary metrics are identifier matching and a count, neither of which leaves the judge much discretion. The five dense with defects documents are one genre, and the claims here extend to critical review of technical engineering documents rather than to prompting in general. The two claude -p transports carry a partial version of the first wrapper's instruction stack that the others do not, which is registered in the design as limitation 5 and is one of the reasons every primary claim stays within model. The contrasts rest on five paired documents at four degrees of freedom, so the intervals are wide by construction and a null on this instrument is a weak instrument as much as a weak effect.

The answer

Neither of these two wrappers earned its tokens on this task. Neither produced a gain distinguishable from noise on either confirmatory model, the combination measurably cost claude-opus-5 yield against the key, and the placebo arm holds the only recall interval on either confirmatory model that excludes zero, which is the metric both wrappers are sold on and neither reached.

Saturation explains part of that and not all of it, and the difference is worth stating because the two confirmatory models are not in the same position. On claude-opus-5 the recall route was closed before the grid ran: pooled control recall is 0.875 and the registered bar asks for 0.125 more, so only a perfect sweep in every one of the fifteen runs could have cleared it. gpt-5.6-sol sits at 0.733, and its disagreement-wrapper recall interval of −0.038 to +0.205 contains the effect the design registered as meaningful. On that model this is an underpowered null, not an occupied ceiling, and five paired documents at four degrees of freedom is the reason.

The finding worth carrying forward is about the instrument rather than the wrappers. The dull scaffolding in the plain control (define the target, fix the output shape, ask for a confidence number, forbid padding, close with a recommendation) buys more yield against the key than a hundred words of instruction to ruminate at length, and the arm that asks for the most output is the arm whose extra output this instrument can score the least. A null result here means no qualifying evidence rather than proof of absence, and the pilot suggests the reason is that there was very little room left to win.

Corrections made before publication

This note was written when the study closed and was staged unpublished under a commission rule that tied a full post to a positive result. When the owner ruled on 2026-08-29 that the negative publishes on its own merits, it was re-checked against the experiment's own artifacts before going out, and seven statements did not survive that check. They are listed here rather than quietly amended, because a note whose subject is measurement discipline does not get to correct itself in silence.

what it said what the artifact says where
365 runs 375 scored runs results/grid.json, n_runs_ok and 375 per_run rows
4,084 items re-checked by the second judge 4,179 items, 51 overturned analysis/JUDGE-AGREEMENT.txt
recall must not have fallen recall must not have fallen by more than 0.05 DESIGN.md condition (a)
naive has the worst precision of any arm on every model tested true on claude-opus-5 and claude-fable-5; on gpt-5.6-sol and gpt-5-6-pro the worst arm is placeboA pooled precision in analysis/NUMBERS.txt
telling a model to work harder does not make it better at finding the defect that is actually there not rated by this design; detection on the two clean documents is a single draw registered as a count, and the only confirmatory model with headroom there runs the other way DESIGN.md §9 and limitation 8
telling a model to work harder roughly doubles the criticism it manufactures the doubling belongs to the arms that name no cognitive move, which run 1.8 to 2.8 times the control's assertion count; the two wrappers run 1.0 to 1.3 times it clean-document assertion counts in analysis/NUMBERS.txt
the two sentences separating naive from control are what buy the precision arms/base.txt and arms/naive.txt differ by five things, varied together, so which of them does the work is not identified the arm prompt files
the two near-clean documents carry "exactly one real defect apiece" in the instrument section too corrected there as well as in the discrimination section, to one registered defect judge/opus-extras/
placeboB absent from the clean-document table both placebos are now reported in every table here and in the post analysis/NUMBERS.txt

A second review round, a four-persona council whose synthesis re-derived each finding against the experiment directory, found four more that a reader deserves to see named:

what it said what the artifact says where
the placebo's fall in precision means it traded accuracy for volume crediting the assertions the second judge sustained, the share of true assertions barely moves (Sol 0.994 to 0.990, Opus 0.857 to 0.853); the fall is the answer key being diluted analysis/VALIDITY.txt
the two near-clean documents carry exactly one real defect each the second judge sustained 1,088 further real weaknesses in them, 19 and 16 distinct passages found in ten or more runs apiece judge/opus-extras/, deduplicated in analysis/VALIDITY.txt
the control bundle buys its yield "at no cost in tokens" it did not raise measured output tokens, which is the only token figure recorded; the control prompt is 89 words against the bare arm's 13 arms/base.txt, arms/naive.txt
both wrappers "quoted here in full" they were excerpted; the first wrapper's closing anti-inflation clause was missing entirely, and both are now printed in full arms/T1.txt, arms/T2.txt

Also corrected in this round: the combination arm's interval was printed with its signs dropped, so the only significant wrapper-direction result in the grid read as a gain; the threshold restatement omitted the recall upper-bound clause and the non-falling clauses; "two permanently missing cells" became two missing runs, since both cells survive at two samples; and the claim that these models arrive at a review task near the ceiling was narrowed to claude-opus-5, because gpt-5.6-sol sits at 0.733 and its disagreement-wrapper interval still contains the registered effect.

One figure is left exactly as the study wrote it, because correcting a study's own number is not a publication decision. This note reports the Opus authorship gap as 0.875 against 0.763, a gap of 0.112, while results/grid.json carries opus_authorship_advantage 0.1139 against a mean of 0.7611. The two trace to a disagreement between two of the study's own artifacts about one figure: analysis/NUMBERS.txt lists gpt-5-6-pro control recall as 0.781, while the per-run mean in results/grid.json is 0.775, and that difference is the whole of the 0.763 against 0.7611 gap. The withdrawal rule fires on either, so nothing downstream moves.

Appendix: every cell

Every combination of model, document and arm appears below, including the cells that are missing and the ones that were never run. Sample count is the number of completed draws. Recall is the fraction of that document's planted defects credited, with the count in brackets; the two clean documents carry one defect each, so recall on them is a single draw and is reported as a count rather than a rate. Yield is credited planted defects over problems asserted, which is what the appendix column labelled yield reports; it is not an accuracy rate, for the reason given above. Cost comes from each transport's own accounting: wall seconds and output tokens for the command line transports, and for gpt-5-6-pro the end to end broker time, which includes queueing and is not model time, together with the size of the delivered answer, because the broker exposes no token count.

No cell in the grid was called off. Every one of the 112 core cells, meaning the plain control and the three wrapper arms on all seven documents across all four models, was run and scored. claude-fable-5 ran all seven arms on all seven documents and is exploratory because it ran at one sample per cell rather than three, not because it was stopped early. gpt-5-6-pro ran the four core arms on all seven documents plus the matched placebo on six; the two arms outside its plan are marked as such rather than left blank. Two gpt-5.6-sol cells are permanently missing, both on the clean script, under the disagreement wrapper and under the two wrappers combined, each rejected twice for reaching outside the document; under the attrition rule set in advance they count as missing rather than as zero, and they are the only absent cells in the grid.

The same figures, with the individual runs behind them, are in results/grid.json.

claude-opus-5

topic arm n recall yield asserted false alarms cost
A naive 3 0.708 (5.7/8) 0.200 29.0 5.00 167 s, 10382 tok
A control 3 0.625 (5.0/8) 0.326 16.0 1.67 139 s, 8943 tok
A T1 3 0.750 (6.0/8) 0.318 19.0 2.67 177 s, 11352 tok
A T2 3 0.708 (5.7/8) 0.413 14.0 1.00 134 s, 8639 tok
A T1T2 3 0.667 (5.3/8) 0.209 25.7 6.00 209 s, 13649 tok
A placeboA 3 0.708 (5.7/8) 0.221 28.0 2.33 222 s, 14842 tok
A placeboB 3 0.750 (6.0/8) 0.257 26.3 4.67 194 s, 12665 tok
B naive 3 0.917 (7.3/8) 0.240 30.7 9.00 173 s, 11895 tok
B control 3 0.917 (7.3/8) 0.423 18.7 2.67 141 s, 9246 tok
B T1 3 0.875 (7.0/8) 0.344 21.7 5.00 194 s, 12727 tok
B T2 3 0.917 (7.3/8) 0.384 20.3 3.00 144 s, 9275 tok
B T1T2 3 0.958 (7.7/8) 0.269 29.0 6.33 202 s, 12868 tok
B placeboA 3 0.917 (7.3/8) 0.225 34.7 9.33 256 s, 17311 tok
B placeboB 3 0.917 (7.3/8) 0.236 31.0 7.33 190 s, 12569 tok
C naive 3 0.917 (7.3/8) 0.237 31.3 5.33 143 s, 9145 tok
C control 3 0.917 (7.3/8) 0.383 20.7 4.00 128 s, 8236 tok
C T1 3 0.875 (7.0/8) 0.217 33.7 4.00 224 s, 14799 tok
C T2 3 0.917 (7.3/8) 0.348 23.0 5.33 147 s, 9643 tok
C T1T2 3 0.917 (7.3/8) 0.216 34.0 6.00 215 s, 14096 tok
C placeboA 3 0.875 (7.0/8) 0.261 27.0 2.33 248 s, 16829 tok
C placeboB 3 0.958 (7.7/8) 0.311 24.7 3.67 198 s, 13207 tok
D naive 3 0.958 (7.7/8) 0.294 26.7 2.33 124 s, 7722 tok
D control 3 0.917 (7.3/8) 0.407 18.7 1.67 96 s, 6055 tok
D T1 3 0.917 (7.3/8) 0.445 17.3 0.67 157 s, 10211 tok
D T2 3 0.917 (7.3/8) 0.511 15.3 0.00 118 s, 7496 tok
D T1T2 3 0.875 (7.0/8) 0.379 19.7 2.00 182 s, 11884 tok
D placeboA 3 0.917 (7.3/8) 0.207 40.0 5.00 234 s, 15870 tok
D placeboB 3 0.875 (7.0/8) 0.251 28.7 1.33 182 s, 12238 tok
E naive 3 0.958 (7.7/8) 0.195 39.3 9.67 150 s, 9483 tok
E control 3 1.000 (8.0/8) 0.298 28.0 7.00 121 s, 7499 tok
E T1 3 1.000 (8.0/8) 0.343 28.0 2.67 234 s, 14903 tok
E T2 3 0.917 (7.3/8) 0.388 20.3 2.33 117 s, 7249 tok
E T1T2 3 1.000 (8.0/8) 0.262 30.7 4.33 200 s, 12569 tok
E placeboA 3 1.000 (8.0/8) 0.259 31.7 6.00 238 s, 15559 tok
E placeboB 3 0.958 (7.7/8) 0.192 43.3 9.67 203 s, 13303 tok
F naive 3 1.000 (1.0/1) 0.040 25.3 10.33 207 s, 13769 tok
F control 3 1.000 (1.0/1) 0.087 11.7 3.00 174 s, 11570 tok
F T1 3 1.000 (1.0/1) 0.076 13.7 2.33 249 s, 16744 tok
F T2 3 1.000 (1.0/1) 0.071 14.7 2.67 184 s, 11835 tok
F T1T2 3 1.000 (1.0/1) 0.058 17.3 7.67 274 s, 18166 tok
F placeboA 3 1.000 (1.0/1) 0.043 26.3 5.00 326 s, 22103 tok
F placeboB 3 1.000 (1.0/1) 0.043 23.3 7.00 276 s, 18590 tok
G naive 3 1.000 (1.0/1) 0.050 20.0 5.33 146 s, 9578 tok
G control 3 1.000 (1.0/1) 0.079 13.7 1.67 155 s, 10492 tok
G T1 3 1.000 (1.0/1) 0.092 12.3 1.67 236 s, 16564 tok
G T2 3 1.000 (1.0/1) 0.068 16.7 3.33 182 s, 12451 tok
G T1T2 3 0.667 (0.7/1) 0.049 14.3 1.33 260 s, 17779 tok
G placeboA 3 1.000 (1.0/1) 0.037 29.7 7.00 270 s, 18649 tok
G placeboB 3 1.000 (1.0/1) 0.036 28.3 7.00 249 s, 17322 tok

gpt-5.6-sol xhigh

topic arm n recall yield asserted false alarms cost
A naive 3 0.542 (4.3/8) 0.292 19.0 2.00 161 s, 21171 tok
A control 3 0.500 (4.0/8) 0.381 10.7 0.00 108 s, 13810 tok
A T1 3 0.667 (5.3/8) 0.421 12.7 0.00 266 s, 25368 tok
A T2 3 0.542 (4.3/8) 0.408 10.7 0.00 146 s, 21168 tok
A T1T2 3 0.583 (4.7/8) 0.424 11.0 0.00 236 s, 19328 tok
A placeboA 3 0.667 (5.3/8) 0.229 24.3 0.00 356 s, 26850 tok
A placeboB 3 0.583 (4.7/8) 0.245 19.0 0.00 272 s, 28238 tok
B naive 3 0.833 (6.7/8) 0.366 18.3 1.67 119 s, 21710 tok
B control 3 0.625 (5.0/8) 0.500 10.7 0.00 94 s, 12920 tok
B T1 3 0.833 (6.7/8) 0.645 10.3 0.00 232 s, 25122 tok
B T2 3 0.708 (5.7/8) 0.496 12.3 0.67 148 s, 20663 tok
B T1T2 3 0.792 (6.3/8) 0.552 12.3 0.00 271 s, 25826 tok
B placeboA 3 0.792 (6.3/8) 0.349 20.0 0.33 299 s, 23286 tok
B placeboB 3 0.792 (6.3/8) 0.465 14.0 0.00 246 s, 19363 tok
C naive 3 0.833 (6.7/8) 0.278 24.7 2.00 99 s, 18196 tok
C control 3 0.792 (6.3/8) 0.530 12.0 0.00 247 s, 15691 tok
C T1 3 0.792 (6.3/8) 0.392 19.3 0.67 219 s, 24786 tok
C T2 3 0.792 (6.3/8) 0.487 13.0 0.00 135 s, 14588 tok
C T1T2 3 0.792 (6.3/8) 0.365 22.0 0.00 196 s, 18869 tok
C placeboA 3 0.833 (6.7/8) 0.246 27.3 0.00 300 s, 22411 tok
C placeboB 3 0.875 (7.0/8) 0.441 16.0 0.00 264 s, 23347 tok
D naive 3 0.875 (7.0/8) 0.502 14.7 0.33 58 s, 16853 tok
D control 3 0.792 (6.3/8) 0.683 9.3 0.00 42 s, 21677 tok
D T1 3 0.833 (6.7/8) 0.468 14.7 0.00 115 s, 19761 tok
D T2 3 0.875 (7.0/8) 0.586 12.0 0.00 72 s, 12120 tok
D T1T2 3 0.792 (6.3/8) 0.513 12.3 0.00 168 s, 14800 tok
D placeboA 3 0.833 (6.7/8) 0.420 16.3 0.00 146 s, 15094 tok
D placeboB 3 0.875 (7.0/8) 0.492 14.3 0.00 85 s, 12916 tok
E naive 3 0.917 (7.3/8) 0.440 16.7 0.33 99 s, 12864 tok
E control 3 0.958 (7.7/8) 0.678 11.3 0.33 97 s, 13263 tok
E T1 3 0.958 (7.7/8) 0.639 13.0 0.33 222 s, 17967 tok
E T2 3 0.958 (7.7/8) 0.616 12.7 0.00 100 s, 13071 tok
E T1T2 3 0.958 (7.7/8) 0.530 15.0 0.00 260 s, 19582 tok
E placeboA 3 1.000 (8.0/8) 0.271 30.7 1.33 289 s, 20287 tok
E placeboB 3 0.917 (7.3/8) 0.435 17.0 0.00 230 s, 17918 tok
F naive 3 0.000 (0.0/1) 0.000 15.7 2.00 224 s, 17585 tok
F control 3 0.000 (0.0/1) 0.000 6.0 0.00 175 s, 15758 tok
F T1 2 0.000 (0.0/1) 0.000 9.0 0.00 277 s, 29519 tok
F T2 3 0.333 (0.3/1) 0.042 7.7 0.00 188 s, 22135 tok
F T1T2 2 0.000 (0.0/1) 0.000 10.5 0.00 238 s, 27830 tok
F placeboA 3 0.000 (0.0/1) 0.000 13.3 0.00 410 s, 25812 tok
F placeboB 3 0.667 (0.7/1) 0.067 10.3 0.00 318 s, 20747 tok
G naive 3 0.667 (0.7/1) 0.065 10.0 0.33 44 s, 10776 tok
G control 3 0.667 (0.7/1) 0.114 6.0 0.00 59 s, 11629 tok
G T1 3 1.000 (1.0/1) 0.145 7.0 0.00 133 s, 15085 tok
G T2 3 1.000 (1.0/1) 0.159 6.3 0.00 82 s, 12552 tok
G T1T2 3 1.000 (1.0/1) 0.132 7.7 0.67 166 s, 16498 tok
G placeboA 3 1.000 (1.0/1) 0.065 16.3 0.00 235 s, 18268 tok
G placeboB 3 1.000 (1.0/1) 0.089 11.3 0.00 133 s, 18167 tok

claude-fable-5

topic arm n recall yield asserted false alarms cost
A naive 1 0.625 (5.0/8) 0.263 19.0 5.00 137 s, 9120 tok
A control 1 0.625 (5.0/8) 0.208 24.0 5.00 100 s, 6642 tok
A T1 1 0.500 (4.0/8) 0.333 12.0 0.00 195 s, 13473 tok
A T2 1 0.500 (4.0/8) 0.210 19.0 2.00 132 s, 8712 tok
A T1T2 1 0.625 (5.0/8) 0.556 9.0 0.00 188 s, 12901 tok
A placeboA 1 0.875 (7.0/8) 0.636 11.0 0.00 186 s, 13071 tok
A placeboB 1 1.000 (8.0/8) 0.286 28.0 5.00 211 s, 14680 tok
B naive 1 0.875 (7.0/8) 0.438 16.0 3.00 128 s, 8396 tok
B control 1 0.750 (6.0/8) 0.400 15.0 3.00 87 s, 5367 tok
B T1 1 1.000 (8.0/8) 0.444 18.0 3.00 157 s, 10722 tok
B T2 1 0.875 (7.0/8) 0.636 11.0 0.00 102 s, 6685 tok
B T1T2 1 1.000 (8.0/8) 0.444 18.0 3.00 216 s, 15325 tok
B placeboA 1 0.750 (6.0/8) 0.207 29.0 6.00 139 s, 9924 tok
B placeboB 1 1.000 (8.0/8) 0.381 21.0 3.00 128 s, 8731 tok
C naive 1 0.750 (6.0/8) 0.214 28.0 7.00 104 s, 6500 tok
C control 1 0.750 (6.0/8) 0.667 9.0 0.00 88 s, 5662 tok
C T1 1 0.750 (6.0/8) 0.231 26.0 8.00 206 s, 14429 tok
C T2 1 0.750 (6.0/8) 0.600 10.0 1.00 130 s, 8830 tok
C T1T2 1 0.875 (7.0/8) 0.292 24.0 6.00 145 s, 9926 tok
C placeboA 1 0.875 (7.0/8) 0.350 20.0 4.00 198 s, 14313 tok
C placeboB 1 1.000 (8.0/8) 0.235 34.0 5.00 162 s, 11402 tok
D naive 1 0.875 (7.0/8) 0.467 15.0 0.00 79 s, 4758 tok
D control 1 0.875 (7.0/8) 0.700 10.0 1.00 58 s, 3284 tok
D T1 1 1.000 (8.0/8) 0.471 17.0 0.00 115 s, 7642 tok
D T2 1 0.625 (5.0/8) 0.714 7.0 1.00 69 s, 4207 tok
D T1T2 1 0.750 (6.0/8) 0.667 9.0 0.00 87 s, 5524 tok
D placeboA 1 0.875 (7.0/8) 0.500 14.0 0.00 154 s, 10551 tok
D placeboB 1 0.875 (7.0/8) 0.583 12.0 2.00 86 s, 5654 tok
E naive 1 1.000 (8.0/8) 0.276 29.0 11.00 94 s, 5748 tok
E control 1 0.875 (7.0/8) 0.500 14.0 2.00 75 s, 4483 tok
E T1 1 0.750 (6.0/8) 0.429 14.0 2.00 124 s, 7952 tok
E T2 1 0.875 (7.0/8) 0.389 18.0 2.00 86 s, 5231 tok
E T1T2 1 1.000 (8.0/8) 0.381 21.0 3.00 127 s, 8327 tok
E placeboA 1 1.000 (8.0/8) 0.421 19.0 3.00 126 s, 8222 tok
E placeboB 1 0.875 (7.0/8) 0.538 13.0 1.00 95 s, 5999 tok
F naive 1 1.000 (1.0/1) 0.067 15.0 4.00 139 s, 9604 tok
F control 1 1.000 (1.0/1) 0.167 6.0 0.00 119 s, 8029 tok
F T1 1 1.000 (1.0/1) 0.091 11.0 3.00 216 s, 15263 tok
F T2 1 1.000 (1.0/1) 0.167 6.0 0.00 123 s, 8453 tok
F T1T2 1 1.000 (1.0/1) 0.200 5.0 1.00 175 s, 12092 tok
F placeboA 1 1.000 (1.0/1) 0.056 18.0 1.00 208 s, 14892 tok
F placeboB 1 1.000 (1.0/1) 0.071 14.0 0.00 210 s, 14459 tok
G naive 1 1.000 (1.0/1) 0.062 16.0 3.00 94 s, 6171 tok
G control 1 1.000 (1.0/1) 0.200 5.0 0.00 90 s, 6358 tok
G T1 1 1.000 (1.0/1) 0.250 4.0 0.00 112 s, 7423 tok
G T2 1 1.000 (1.0/1) 0.143 7.0 0.00 112 s, 7465 tok
G T1T2 1 1.000 (1.0/1) 0.143 7.0 0.00 92 s, 6168 tok
G placeboA 1 1.000 (1.0/1) 0.100 10.0 2.00 137 s, 9754 tok
G placeboB 1 1.000 (1.0/1) 0.143 7.0 2.00 117 s, 8140 tok

gpt-5-6-pro

topic arm n recall yield asserted false alarms cost
A naive 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
A control 1 0.500 (4.0/8) 0.400 10.0 0.00 1230 s broker, 10172 chars
A T1 1 0.500 (4.0/8) 0.333 12.0 0.00 1230 s broker, 28755 chars
A T2 1 0.625 (5.0/8) 0.455 11.0 0.00 1776 s broker, 14880 chars
A T1T2 1 0.500 (4.0/8) 0.286 14.0 0.00 1308 s broker, 34649 chars
A placeboA 1 0.875 (7.0/8) 0.226 31.0 0.00 2346 s broker, 24094 chars
A placeboB 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
B naive 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
B control 1 0.750 (6.0/8) 0.600 10.0 1.00 1176 s broker, 7090 chars
B T1 1 0.875 (7.0/8) 0.200 35.0 4.00 2454 s broker, 20741 chars
B T2 1 0.750 (6.0/8) 0.250 24.0 2.00 1080 s broker, 10335 chars
B T1T2 1 0.875 (7.0/8) 0.500 14.0 1.00 3600 s broker, 29123 chars
B placeboA 1 0.875 (7.0/8) 0.412 17.0 0.00 121980 s broker, 16636 chars
B placeboB 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
C naive 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
C control 1 0.750 (6.0/8) 0.462 13.0 0.00 1884 s broker, 9938 chars
C T1 1 0.875 (7.0/8) 0.241 29.0 0.00 2454 s broker, 20466 chars
C T2 1 0.875 (7.0/8) 0.583 12.0 0.00 2448 s broker, 10710 chars
C T1T2 1 0.750 (6.0/8) 0.353 17.0 0.00 3024 s broker, 27942 chars
C placeboA 0 not run not run not run not run not run
C placeboB 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
D naive 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
D control 1 0.875 (7.0/8) 0.700 10.0 0.00 1086 s broker, 8784 chars
D T1 1 0.875 (7.0/8) 0.778 9.0 0.00 1164 s broker, 14935 chars
D T2 1 0.875 (7.0/8) 0.304 23.0 0.00 121974 s broker, 11321 chars
D T1T2 1 0.875 (7.0/8) 0.333 21.0 0.00 3600 s broker, 22813 chars
D placeboA 1 0.875 (7.0/8) 0.333 21.0 0.00 121974 s broker, 17071 chars
D placeboB 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
E naive 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
E control 1 1.000 (8.0/8) 0.667 12.0 0.00 2346 s broker, 11389 chars
E T1 1 1.000 (8.0/8) 0.615 13.0 0.00 2346 s broker, 18359 chars
E T2 1 0.875 (7.0/8) 0.636 11.0 0.00 2940 s broker, 15081 chars
E T1T2 1 0.875 (7.0/8) 0.700 10.0 0.00 1308 s broker, 20407 chars
E placeboA 1 1.000 (8.0/8) 0.267 30.0 1.00 3558 s broker, 24925 chars
E placeboB 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
F naive 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
F control 1 0.000 (0.0/1) 0.000 7.0 0.00 1206 s broker, 6441 chars
F T1 1 0.000 (0.0/1) 0.000 10.0 1.00 2358 s broker, 17991 chars
F T2 1 0.000 (0.0/1) 0.000 11.0 1.00 2358 s broker, 9939 chars
F T1T2 1 0.000 (0.0/1) 0.000 11.0 1.00 4080 s broker, 20708 chars
F placeboA 1 1.000 (1.0/1) 0.062 16.0 2.00 2454 s broker, 16680 chars
F placeboB 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
G naive 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)
G control 1 1.000 (1.0/1) 0.167 6.0 0.00 1776 s broker, 6250 chars
G T1 1 1.000 (1.0/1) 0.143 7.0 0.00 2958 s broker, 18765 chars
G T2 1 1.000 (1.0/1) 0.143 7.0 0.00 2952 s broker, 8729 chars
G T1T2 1 1.000 (1.0/1) 0.125 8.0 0.00 4080 s broker, 16023 chars
G placeboA 1 1.000 (1.0/1) 0.062 16.0 0.00 2448 s broker, 15667 chars
G placeboB 0 not run not run not run not run arm outside the Pro plan (four core arms plus the matched placebo)