Two prompt wrappers circulated on 2026-08-27, both claiming to raise the quality of a model's work without changing the task it is given, both of them prefixes to an otherwise untouched prompt. The first, whose author claimed it "works incredibly well with GPT-5.6 Sol xhigh", asks the model to "meticulously and comprehensively study" a document, to say whether it sees "anything you disagree with… anything that is obviously contradictory", to "meditate and ruminate on all of it deeply and profoundly", and finally to report "how confident you REALLY are". The second, whose author claimed it "transforms base metals into gold", instructs the model that "for every decision, ask what the best expert in that field would do and why they would reject your current choice", and that "every trade-off you take must be stated to the user, never absorbed".
The excerpts above are excerpts. Both wrappers were tested verbatim, prepended to an otherwise identical task prompt, and both are printed in full below, because one of them ends on a clause that changes how it should be read. Neither author is named here; each is referred to only as the author of the wrapper being tested.
The first wrapper, in full:
I want you to meticulously and comprehensively study
; do you see anything you disagree with? Anything that is obviously contradictory? Foolishly inconsistent? Misguided and unlikely to work in practice? Ask yourself WHY you feel that way and meditate and ruminate on all of it deeply and profoundly. Finally, explain your thought process and conclusions and your justifications/rationale for them, as well as your confidence in those (and also how you know you're confident, and how confident you REALLY are!). Important: don't disagree just to be disagreeable, it's OK to just agree if you think it's right!
That closing sentence is an anti-inflation clause, so this wrapper does not simply ask for more criticism, and a reading of the result that treats it as an instruction to find more is reading a wrapper it did not test.
The second wrapper, in full:
For every decision, ask what the best expert in that field would do and why they would reject your current choice; if you can name that reason, don't make the choice. Optimize for what that expert would judge correct, never for what satisfies the stated constraints most cheaply. Every trade-off you take must be stated to the user, never absorbed.
The question a test like this has to answer is not whether a wrapper improves on nothing, since almost any instruction improves on nothing. It is whether the specific cognitive move each wrapper names does more than an instruction of the same length that names no move at all. That control is the whole design, and it is what produced the result.
The instrument, and the finding that came before the experiment
Seven documents were written for this: five engineering artefacts carrying eight planted defects each (a restart design whose central concurrency claim is false as implemented, a watcher script with reliability defects drawn from a real taxonomy, a database migration plan that contradicts its own invariants, a results report with an arithmetic error and a confounded causal claim, and a decision memo recommending a job queue that loses jobs), plus two documents written to be correct, each carrying exactly one registered defect. That registration turned out not to describe them, which is the subject of a later section. Each defect has a narrow credit rule stating what counts as having found it, and defects are atomic, so one asserted problem earns at most one identifier.
A calibration pilot ran the plain control prompt once per document before any wrapper was tested, and it returned mean recall of 0.80. The documents were then rewritten longer, denser and subtler, with eight defects each rather than five. Recall came back at 0.80 again.
That number is the finding underneath every other finding in this note. On technical documents dense with defects, claude-opus-5 and gpt-5.6-sol at their top reasoning settings are already close to
saturated at a plain task prompt, which leaves very little room for any wrapper to occupy. Only
the concurrency design sat low enough to move (pooled control recall 0.562, which is 0.62 on
claude-opus-5 and 0.50 on gpt-5.6-sol read within model); the decision memo sat at
0.969, one defect from the ceiling. An experiment resting on recall alone would have measured a quantity
that could not change, so yield against the key and behaviour on the two clean documents were promoted
alongside recall before the grid ran.
The grid
Seven arms ran on every document: the two wrappers, their combination, a plain control, a bare
naive prompt carrying no request for confidence and no instruction against padding, and two
placebos matched in length and imperative density to the two wrappers. The placebos push breadth
and effort while naming no cognitive move; placeboA says, in full, to work carefully and
completely, to aim for comprehensive coverage, to list every problem found rather than only the
obvious ones, and not to stop early.
The confirmatory grid is claude-opus-5 and gpt-5.6-sol at xhigh, targeting three samples per
cell and returning 292 of a planned 294 runs, with
contrasts computed as paired differences across documents. claude-fable-5 and 39
gpt-5-6-pro dives ran as exploratory arms and decide nothing. Models were driven only on subscription routes, with no API key at any vendor. claude-opus-5 and
claude-fable-5 ran through claude -p --model opus|fable --effort xhigh with every tool disabled;
gpt-5.6-sol ran at xhigh through codex exec -m gpt-5.6-sol, in a fresh empty directory per run,
with any transcript showing a tool call rejected; gpt-5-6-pro ran through a request broker that
serialises dives across the machine.
Subjects had no tools: a canary test during design proved that a default run reads files on disk, which would have put the answer key within reach, so file access was disabled and any run whose transcript shows a tool call is rejected.
The threshold, stated before the data
Before the threshold, the second measure needs its right name. planted_precision is credited
planted defects divided by every assertion a run makes, so an observation that is true of the
document but absent from the answer key is charged against it exactly as a fabrication is. It is a
yield against the key, not an accuracy rate, and this note calls it that wherever the distinction
carries weight. What separates the two is measured further down.
A wrapper counted as having worked, for a given model, only if all five of the following held. Recall must not have fallen by more than 0.05, with the upper bound of its interval against control at or above zero. There must be a gain on a measure with room to move, meaning either recall up by at least 0.125 (one full defect of eight) with yield not falling, or yield up by at least 0.10 with recall not falling. The interval on whichever measure carried that gain must exclude zero. The same gain must hold against the placebo matched to that wrapper, at half the margin or better, with its interval also excluding zero. And assertions against the two near-clean documents must not rise by more than 25 per cent. A verdict of yes required the same wrapper to clear all five on both confirmatory models.
Two things about that threshold are worth stating plainly, because neither flatters the null. The
matched-placebo condition is defined for T1 and T2; the combination arm is 157 words and the design
maps it to the 91-word placeboA, so it has no comparator of its own. And the intervals throughout
are paired by document across the five detection topics, five pairs and four degrees of freedom,
Student t at 95 per cent, which is a small instrument to hang a null on.
None of the three wrapper arms cleared even the second condition on either model. The commission that ordered the test tied a full post to a yes, so none was drafted when the grid closed; the owner ruled on 2026-08-29 that a decision-relevant negative publishes on its own merits rather than on its sign, and the scoped post this note accompanies is the result. A no here means no qualifying evidence rather than demonstrated absence, since the design carries no equivalence test.
No wrapper passed
Every recall interval for the two wrappers and their combination includes zero, on both confirmatory
models. The single interval in the grid that excludes zero in a wrapper's direction is a harm: the
combination costs claude-opus-5 0.100 of yield, with an interval running from −0.181 to
−0.019.
| arm | Opus recall Δ | Opus yield Δ | Sol recall Δ | Sol yield Δ |
|---|---|---|---|---|
| T1 (meticulous disagreement) | +0.008 | −0.034 | +0.083 | −0.041 |
| T2 (best expert would reject) | +0.000 | +0.041 | +0.042 | −0.035 |
| T1 and T2 together | +0.008 | −0.100 | +0.050 | −0.077 |
| placeboA (try harder) | +0.008 | −0.133 | +0.092 | −0.251 |
| placeboB (try harder, shorter) | +0.017 | −0.118 | +0.075 | −0.139 |
| naive (bare instruction) | +0.017 | −0.134 | +0.067 | −0.179 |
Bold marks an interval excluding zero. All figures are differences against the plain control, paired by document across the five documents carrying eight defects each.
The placebo produced the only recall gain in the grid that is distinguishable from noise. On
gpt-5.6-sol, placeboA gains 0.092 with an interval from 0.007 to 0.177, which neither wrapper
achieved on either model. It is one nominal interval among twelve contrasts, uncorrected, and the
design's declared defence against multiplicity is that a yes requires the same wrapper on the same
metric route on both confirmatory models. This result does not have that, so it is a signal to
follow rather than a finding to bank. On the concurrency design, the one document with real
headroom, the placebos reach 0.88 and 1.00 on claude-fable-5 and 0.88 on gpt-5-6-pro, where the
disagreement wrapper reaches 0.50 on both.
What the placebo buys is volume. It adds 12.9 asserted problems per run on Sol, and its fall in yield against the answer key, 0.251, is the largest in the grid. Recall counts how many planted defects appear somewhere in a list, so a longer list is rewarded mechanically, and an instruction to list everything produces a longer list. That mechanism is available to anyone willing to write one sentence.
What it does not buy is accuracy either, and the difference between those two statements is the
second judge's work. Crediting the assertions that judge sustained as real weaknesses alongside the
planted hits, the share of a run's assertions that are true barely moves: on gpt-5.6-sol the
control sits at 0.994 and placeboA at 0.990, and on claude-opus-5 the control sits at 0.857 and
placeboA at 0.853. Adjudicated false alarms per run go from 0.07 to 0.33 on Sol and from 3.40 to
5.00 on Opus. The collapse in yield is therefore the answer key being diluted rather than the model
being wrong a larger share of the time. The sustained share does slip, from 0.994 to 0.990 on Sol,
which moves the adjudicated-wrong share from 0.006 to 0.010 off a very small base, and the design
carries no test on a movement that size; what it does not do is rise. That is a narrower and less
flattering result than "the placebo degrades the answers": what it degrades is the fraction of its
output that this instrument can score.
The discrimination endpoint, and why it cannot be read as written
Two documents were written to be correct, each carrying a single registered weakness: a watcher script that handles its own error paths properly but goes silent when its endpoint fails, and a report whose statistics are honest but whose recommendation extrapolates to two regions it has no data for. The design's condition (e) asks that assertions against them not rise by more than 25 per cent, and the reading originally written here was that an inflated count meant manufactured criticism against a document that is essentially fine.
The study's own second judge does not support that reading. Across those two documents it sustained 1,088 assertions as real weaknesses the answer key never listed, and deduplicating them by the passage each quotes leaves 60 distinct spans in the script and 96 in the report; of those, 19 and 16 respectively survive a further filter for having been found in ten or more independent runs, which is a replication threshold rather than part of the deduplication. The two documents are not near-clean. The control arm is already finding their unregistered weaknesses at a rate that makes the assertion count uninterpretable on its own.
| model | arm | asserted | sustained as real | adjudicated false alarms |
|---|---|---|---|---|
claude-opus-5 |
control | 12.7 | 9.2 | 2.33 |
claude-opus-5 |
T1 | 13.0 | 10.0 | 2.00 |
claude-opus-5 |
T2 | 15.7 | 11.5 | 3.00 |
claude-opus-5 |
T1T2 | 15.8 | 10.5 | 4.50 |
claude-opus-5 |
placeboA | 28.0 | 20.5 | 6.00 |
claude-opus-5 |
placeboB | 25.8 | 17.8 | 7.00 |
claude-opus-5 |
naive | 22.7 | 13.7 | 7.83 |
gpt-5.6-sol |
control | 6.0 | 5.2 | 0.00 |
gpt-5.6-sol |
T1 | 7.8 | 7.2 | 0.00 |
gpt-5.6-sol |
T2 | 7.0 | 6.0 | 0.00 |
gpt-5.6-sol |
T1T2 | 8.8 | 7.8 | 0.40 |
gpt-5.6-sol |
placeboA | 14.8 | 13.3 | 0.00 |
gpt-5.6-sol |
placeboB | 10.8 | 9.5 | 0.00 |
gpt-5.6-sol |
naive | 12.8 | 11.2 | 1.17 |
Per run, on the two documents registered as carrying one defect each. The three columns do not sum
to the asserted count. The remainder runs between 0.50 and 1.50 assertions per run and is the
registered defect itself where an arm found it, plus the small share of extras that judge overturned. Most of what the count picks up is sustained by that judge, so the
count measures how many weaknesses it will sustain rather than how much criticism an arm invents. What the last column does not track is effort language, since the highest false-alarm rate on both
confirmatory models belongs to the bare prompt, which contains none:
claude-opus-5 runs 2.33 false alarms per run at the control, 6.00 under placeboA and 7.83 under
the bare prompt, while on gpt-5.6-sol the column is zero everywhere except 1.17 under the bare
prompt and 0.40 under T1T2. Whether any arm is better at finding the one registered weakness is a
question the design declined to rate, since recall there is a single draw registered as a count
rather than a rate.
Two of the wrapper arms breach condition (e) on gpt-5.6-sol even so, at 7.8 and 8.8 assertions
against the control's 6.0, which is 30 and 47 per cent. Both arms had already failed on the gain
condition, so nothing in the verdict moves; it is recorded because a threshold registered in advance
should be reported when it is crossed and not only when it is convenient.
On gpt-5.6-sol the disagreement wrapper also did something stranger: told to find things it
disagreed with in a document with almost nothing wrong, it reached outside the document for tools,
twice, and both runs were rejected by the hermeticity guard and could not be recovered.
The expert wrapper reduces output, which the instrument scores as nothing
T2 is the only arm anywhere that makes models say less. It cuts Opus assertions from 20.4 to 18.6
and false alarms from 3.40 to 2.33, and it is the only arm that moves Opus yield upward at all,
though that interval still includes zero. A gpt-5-6-pro answer stated the mechanism while
performing it: "To avoid padding, I combined the closely coupled crash-recovery, persistence, and
orphan-cleanup defects into one item."
Consolidation is a real behavioural change, and under atomic scoring it costs a point, since merging two credited defects into one finding surrenders one of them. The measurement is therefore unfair to this particular move. The instrument is not blind to the change, since both the assertion count and the false-alarm count move for T2; what it cannot credit is consolidation on the recall measure. The honest statement is that T2 changes how models write without measurably changing how much of the answer key they find.
The plain clauses beat both wrappers
The naive arm carries a single sentence asking what problems the reader finds. It has the lowest
yield against the key of any arm on claude-opus-5 at 0.233 and on claude-fable-5 at 0.331; on
gpt-5.6-sol and gpt-5-6-pro the lowest is placeboA, at 0.303 and 0.309 against naive's 0.375
and its absence from the Pro plan. What holds on both confirmatory models is the contrast that
matters here: naive loses 0.134 of yield against the control on Opus, at 0.233 against 0.367, and
0.179 on Sol, at 0.375 against 0.554, and both are larger than any yield gain a wrapper managed over
the control on either model. It also carries the highest adjudicated false-alarm rate of any arm on
both, at 6.27 per run on Opus and 1.27 on Sol, so unlike the placebo its low yield is not only
dilution.
What separates naive from control is a bundle of five things, not one clause: a definition of
what counts as a substantive problem, a required shape for each finding, a request for a confidence
percentage, an instruction against padding, and a closing one-line recommendation. The bundle buys
more yield against the key than either wrapper does, and it did not increase measured output tokens,
which is the only token figure this study recorded; the control prompt is itself 89 words against the
bare arm's 13, so it is not free on input. Which of its five parts does the work is not identified by
this design, which varied them together. The plain control also already contains the
confidence-elicitation clause the first wrapper asks for, so that arm is a smaller manipulation than
its word count suggests. The wrappers are not free by comparison, since the
disagreement wrapper costs Opus 60 per cent more output tokens and 58 per cent more wall time for no
measured gain, and placeboA costs 101 per cent more tokens and 92 per cent more time.
Two rules that fired against the result
The design registered a withdrawal rule in advance, because the documents and their planted defects were
written by claude-opus-5. That model is also one of the four under test, so if its control recall
exceeded the other models' by more than 0.10, every claim comparing one model with another was to be withdrawn. Opus
scored 0.875 against 0.763 for the other three, a gap of 0.112, so the rule fired and no comparison
across models is made here. Only contrasts taken inside a single model stand, and those were named
primary in the design from the beginning.
The hermeticity guard rejected five Sol runs for tool use and two proved unrecoverable, leaving two
permanently missing runs; both cells survive at two samples rather than three. Five GPT Pro dives were lost to a broker outage, attested structurally as
having produced no answer at all rather than having lost one, and filed again; all eventually landed.
Coverage on gpt-5-6-pro reached six of seven documents at four arms, the seventh dropped under the
degradation order set in the design.
Whether the scoring holds
Two independent checks on the judging ran after the grid, both scored by claude-opus-5 against
gpt-5.6-sol's verdicts.
The first covers every problem the primary judge accepted as a real weakness of a document without its being on the defect list. That category was the one the design flagged as an escape hatch, since a wrapper that produces more persuasive but still wrong criticism would have its costs booked as a bonus. All 4,179 such items were re-checked by asking of each one whether it is true of the document as written, with a verbatim span from the document required to sustain a yes, and 51 were overturned, a rate of 1.2 per cent. The rate is flat across arms, running from 0.9 per cent under the shorter placebo to 1.7 per cent under the bare prompt, so the category is not absorbing a differential cost. Effort was set to medium for this mechanical check after it returned verdicts identical to the top setting on 42 of 42 items across three sampled runs, at roughly half the time.
The second re-scores a quarter of all runs, stratified by model and by arm, on the primary measure. Per-defect agreement between the two judges is 0.963, and mean recall differs by 0.003 between them (0.804 against 0.807). Agreement by arm runs from 0.919 to 0.990, above the 0.80 floor the design set for reporting the measure without a caveat, and the spread across arms is narrow enough that no contrast rests on one arm being scored more generously than another.
Caveats
The result covers one turn only, and the expert wrapper's author claims it specifically for repeated
injection on each message, which this design cannot reach. Both judges are models that also
appear as arms. Answers in free text were mechanically normalised into finding lists before scoring, so
that a verbose finding could not buy more matching surface than a terse one, and the primary metrics
are identifier matching and a count, neither of which leaves the judge much discretion. The five
dense with defects documents are one genre, and the claims here extend to critical review of technical
engineering documents rather than to prompting in general. The two claude -p transports carry a
partial version of the first wrapper's instruction stack that the others do not, which is registered
in the design as limitation 5 and is one of the reasons every primary claim stays within model. The
contrasts rest on five paired documents at four degrees of freedom, so the intervals are wide by
construction and a null on this instrument is a weak instrument as much as a weak effect.
The answer
Neither of these two wrappers earned its tokens on this task. Neither produced a gain
distinguishable from noise on either confirmatory model, the combination measurably cost
claude-opus-5 yield against the key, and the placebo arm holds the only recall interval on either
confirmatory model that excludes zero, which is the metric both wrappers are sold on and neither
reached.
Saturation explains part of that and not all of it, and the difference is worth stating because the
two confirmatory models are not in the same position. On claude-opus-5 the recall route was closed
before the grid ran: pooled control recall is 0.875 and the registered bar asks for 0.125 more, so
only a perfect sweep in every one of the fifteen runs could have cleared it. gpt-5.6-sol sits at
0.733, and its disagreement-wrapper recall interval of −0.038 to +0.205 contains the effect the
design registered as meaningful. On that model this is an underpowered null, not an occupied
ceiling, and five paired documents at four degrees of freedom is the reason.
The finding worth carrying forward is about the instrument rather than the wrappers. The dull scaffolding in the plain control (define the target, fix the output shape, ask for a confidence number, forbid padding, close with a recommendation) buys more yield against the key than a hundred words of instruction to ruminate at length, and the arm that asks for the most output is the arm whose extra output this instrument can score the least. A null result here means no qualifying evidence rather than proof of absence, and the pilot suggests the reason is that there was very little room left to win.
Corrections made before publication
This note was written when the study closed and was staged unpublished under a commission rule that tied a full post to a positive result. When the owner ruled on 2026-08-29 that the negative publishes on its own merits, it was re-checked against the experiment's own artifacts before going out, and seven statements did not survive that check. They are listed here rather than quietly amended, because a note whose subject is measurement discipline does not get to correct itself in silence.
| what it said | what the artifact says | where |
|---|---|---|
| 365 runs | 375 scored runs | results/grid.json, n_runs_ok and 375 per_run rows |
| 4,084 items re-checked by the second judge | 4,179 items, 51 overturned | analysis/JUDGE-AGREEMENT.txt |
| recall must not have fallen | recall must not have fallen by more than 0.05 | DESIGN.md condition (a) |
naive has the worst precision of any arm on every model tested |
true on claude-opus-5 and claude-fable-5; on gpt-5.6-sol and gpt-5-6-pro the worst arm is placeboA |
pooled precision in analysis/NUMBERS.txt |
| telling a model to work harder does not make it better at finding the defect that is actually there | not rated by this design; detection on the two clean documents is a single draw registered as a count, and the only confirmatory model with headroom there runs the other way | DESIGN.md §9 and limitation 8 |
| telling a model to work harder roughly doubles the criticism it manufactures | the doubling belongs to the arms that name no cognitive move, which run 1.8 to 2.8 times the control's assertion count; the two wrappers run 1.0 to 1.3 times it | clean-document assertion counts in analysis/NUMBERS.txt |
the two sentences separating naive from control are what buy the precision |
arms/base.txt and arms/naive.txt differ by five things, varied together, so which of them does the work is not identified |
the arm prompt files |
| the two near-clean documents carry "exactly one real defect apiece" in the instrument section too | corrected there as well as in the discrimination section, to one registered defect | judge/opus-extras/ |
placeboB absent from the clean-document table |
both placebos are now reported in every table here and in the post | analysis/NUMBERS.txt |
A second review round, a four-persona council whose synthesis re-derived each finding against the experiment directory, found four more that a reader deserves to see named:
| what it said | what the artifact says | where |
|---|---|---|
| the placebo's fall in precision means it traded accuracy for volume | crediting the assertions the second judge sustained, the share of true assertions barely moves (Sol 0.994 to 0.990, Opus 0.857 to 0.853); the fall is the answer key being diluted | analysis/VALIDITY.txt |
| the two near-clean documents carry exactly one real defect each | the second judge sustained 1,088 further real weaknesses in them, 19 and 16 distinct passages found in ten or more runs apiece | judge/opus-extras/, deduplicated in analysis/VALIDITY.txt |
| the control bundle buys its yield "at no cost in tokens" | it did not raise measured output tokens, which is the only token figure recorded; the control prompt is 89 words against the bare arm's 13 | arms/base.txt, arms/naive.txt |
| both wrappers "quoted here in full" | they were excerpted; the first wrapper's closing anti-inflation clause was missing entirely, and both are now printed in full | arms/T1.txt, arms/T2.txt |
Also corrected in this round: the combination arm's interval was printed with its signs dropped, so
the only significant wrapper-direction result in the grid read as a gain; the threshold restatement
omitted the recall upper-bound clause and the non-falling clauses; "two permanently missing cells"
became two missing runs, since both cells survive at two samples; and the claim that these models
arrive at a review task near the ceiling was narrowed to claude-opus-5, because gpt-5.6-sol sits
at 0.733 and its disagreement-wrapper interval still contains the registered effect.
One figure is left exactly as the study wrote it, because correcting a study's own number is not a
publication decision. This note reports the Opus authorship gap as 0.875 against 0.763, a gap of
0.112, while results/grid.json carries opus_authorship_advantage 0.1139 against a mean of
0.7611. The two trace to a disagreement between two of the study's own artifacts about one figure:
analysis/NUMBERS.txt lists gpt-5-6-pro control recall as 0.781, while the per-run mean in
results/grid.json is 0.775, and that difference is the whole of the 0.763 against 0.7611 gap. The
withdrawal rule fires on either, so nothing downstream moves.
Appendix: every cell
Every combination of model, document and arm appears below, including the cells that are missing and
the ones that were never run. Sample count is the number of completed draws. Recall is the fraction
of that document's planted defects credited, with the count in brackets; the two clean documents
carry one defect each, so recall on them is a single draw and is reported as a count rather than a
rate. Yield is credited planted defects over problems asserted, which is what the appendix column
labelled yield reports; it is not an accuracy rate, for the reason given above. Cost comes from each transport's own
accounting: wall seconds and output tokens for the command line transports, and for gpt-5-6-pro the
end to end broker time, which includes queueing and is not model time, together with the size of the
delivered answer, because the broker exposes no token count.
No cell in the grid was called off. Every one of the 112 core cells, meaning the plain control and
the three wrapper arms on all seven documents across all four models, was run and scored.
claude-fable-5 ran all seven arms on all seven documents and is exploratory because it ran at one
sample per cell rather than three, not because it was stopped early. gpt-5-6-pro ran the four core
arms on all seven documents plus the matched placebo on six; the two arms outside its plan are marked
as such rather than left blank. Two gpt-5.6-sol cells are permanently missing, both on the clean
script, under the disagreement wrapper and under the two wrappers combined, each rejected twice for
reaching outside the document; under the attrition rule set in advance they count as missing rather
than as zero, and they are the only absent cells in the grid.
The same figures, with the individual runs behind them, are in results/grid.json.
claude-opus-5
| topic | arm | n | recall | yield | asserted | false alarms | cost |
|---|---|---|---|---|---|---|---|
| A | naive |
3 | 0.708 (5.7/8) | 0.200 | 29.0 | 5.00 | 167 s, 10382 tok |
| A | control |
3 | 0.625 (5.0/8) | 0.326 | 16.0 | 1.67 | 139 s, 8943 tok |
| A | T1 |
3 | 0.750 (6.0/8) | 0.318 | 19.0 | 2.67 | 177 s, 11352 tok |
| A | T2 |
3 | 0.708 (5.7/8) | 0.413 | 14.0 | 1.00 | 134 s, 8639 tok |
| A | T1T2 |
3 | 0.667 (5.3/8) | 0.209 | 25.7 | 6.00 | 209 s, 13649 tok |
| A | placeboA |
3 | 0.708 (5.7/8) | 0.221 | 28.0 | 2.33 | 222 s, 14842 tok |
| A | placeboB |
3 | 0.750 (6.0/8) | 0.257 | 26.3 | 4.67 | 194 s, 12665 tok |
| B | naive |
3 | 0.917 (7.3/8) | 0.240 | 30.7 | 9.00 | 173 s, 11895 tok |
| B | control |
3 | 0.917 (7.3/8) | 0.423 | 18.7 | 2.67 | 141 s, 9246 tok |
| B | T1 |
3 | 0.875 (7.0/8) | 0.344 | 21.7 | 5.00 | 194 s, 12727 tok |
| B | T2 |
3 | 0.917 (7.3/8) | 0.384 | 20.3 | 3.00 | 144 s, 9275 tok |
| B | T1T2 |
3 | 0.958 (7.7/8) | 0.269 | 29.0 | 6.33 | 202 s, 12868 tok |
| B | placeboA |
3 | 0.917 (7.3/8) | 0.225 | 34.7 | 9.33 | 256 s, 17311 tok |
| B | placeboB |
3 | 0.917 (7.3/8) | 0.236 | 31.0 | 7.33 | 190 s, 12569 tok |
| C | naive |
3 | 0.917 (7.3/8) | 0.237 | 31.3 | 5.33 | 143 s, 9145 tok |
| C | control |
3 | 0.917 (7.3/8) | 0.383 | 20.7 | 4.00 | 128 s, 8236 tok |
| C | T1 |
3 | 0.875 (7.0/8) | 0.217 | 33.7 | 4.00 | 224 s, 14799 tok |
| C | T2 |
3 | 0.917 (7.3/8) | 0.348 | 23.0 | 5.33 | 147 s, 9643 tok |
| C | T1T2 |
3 | 0.917 (7.3/8) | 0.216 | 34.0 | 6.00 | 215 s, 14096 tok |
| C | placeboA |
3 | 0.875 (7.0/8) | 0.261 | 27.0 | 2.33 | 248 s, 16829 tok |
| C | placeboB |
3 | 0.958 (7.7/8) | 0.311 | 24.7 | 3.67 | 198 s, 13207 tok |
| D | naive |
3 | 0.958 (7.7/8) | 0.294 | 26.7 | 2.33 | 124 s, 7722 tok |
| D | control |
3 | 0.917 (7.3/8) | 0.407 | 18.7 | 1.67 | 96 s, 6055 tok |
| D | T1 |
3 | 0.917 (7.3/8) | 0.445 | 17.3 | 0.67 | 157 s, 10211 tok |
| D | T2 |
3 | 0.917 (7.3/8) | 0.511 | 15.3 | 0.00 | 118 s, 7496 tok |
| D | T1T2 |
3 | 0.875 (7.0/8) | 0.379 | 19.7 | 2.00 | 182 s, 11884 tok |
| D | placeboA |
3 | 0.917 (7.3/8) | 0.207 | 40.0 | 5.00 | 234 s, 15870 tok |
| D | placeboB |
3 | 0.875 (7.0/8) | 0.251 | 28.7 | 1.33 | 182 s, 12238 tok |
| E | naive |
3 | 0.958 (7.7/8) | 0.195 | 39.3 | 9.67 | 150 s, 9483 tok |
| E | control |
3 | 1.000 (8.0/8) | 0.298 | 28.0 | 7.00 | 121 s, 7499 tok |
| E | T1 |
3 | 1.000 (8.0/8) | 0.343 | 28.0 | 2.67 | 234 s, 14903 tok |
| E | T2 |
3 | 0.917 (7.3/8) | 0.388 | 20.3 | 2.33 | 117 s, 7249 tok |
| E | T1T2 |
3 | 1.000 (8.0/8) | 0.262 | 30.7 | 4.33 | 200 s, 12569 tok |
| E | placeboA |
3 | 1.000 (8.0/8) | 0.259 | 31.7 | 6.00 | 238 s, 15559 tok |
| E | placeboB |
3 | 0.958 (7.7/8) | 0.192 | 43.3 | 9.67 | 203 s, 13303 tok |
| F | naive |
3 | 1.000 (1.0/1) | 0.040 | 25.3 | 10.33 | 207 s, 13769 tok |
| F | control |
3 | 1.000 (1.0/1) | 0.087 | 11.7 | 3.00 | 174 s, 11570 tok |
| F | T1 |
3 | 1.000 (1.0/1) | 0.076 | 13.7 | 2.33 | 249 s, 16744 tok |
| F | T2 |
3 | 1.000 (1.0/1) | 0.071 | 14.7 | 2.67 | 184 s, 11835 tok |
| F | T1T2 |
3 | 1.000 (1.0/1) | 0.058 | 17.3 | 7.67 | 274 s, 18166 tok |
| F | placeboA |
3 | 1.000 (1.0/1) | 0.043 | 26.3 | 5.00 | 326 s, 22103 tok |
| F | placeboB |
3 | 1.000 (1.0/1) | 0.043 | 23.3 | 7.00 | 276 s, 18590 tok |
| G | naive |
3 | 1.000 (1.0/1) | 0.050 | 20.0 | 5.33 | 146 s, 9578 tok |
| G | control |
3 | 1.000 (1.0/1) | 0.079 | 13.7 | 1.67 | 155 s, 10492 tok |
| G | T1 |
3 | 1.000 (1.0/1) | 0.092 | 12.3 | 1.67 | 236 s, 16564 tok |
| G | T2 |
3 | 1.000 (1.0/1) | 0.068 | 16.7 | 3.33 | 182 s, 12451 tok |
| G | T1T2 |
3 | 0.667 (0.7/1) | 0.049 | 14.3 | 1.33 | 260 s, 17779 tok |
| G | placeboA |
3 | 1.000 (1.0/1) | 0.037 | 29.7 | 7.00 | 270 s, 18649 tok |
| G | placeboB |
3 | 1.000 (1.0/1) | 0.036 | 28.3 | 7.00 | 249 s, 17322 tok |
gpt-5.6-sol xhigh
| topic | arm | n | recall | yield | asserted | false alarms | cost |
|---|---|---|---|---|---|---|---|
| A | naive |
3 | 0.542 (4.3/8) | 0.292 | 19.0 | 2.00 | 161 s, 21171 tok |
| A | control |
3 | 0.500 (4.0/8) | 0.381 | 10.7 | 0.00 | 108 s, 13810 tok |
| A | T1 |
3 | 0.667 (5.3/8) | 0.421 | 12.7 | 0.00 | 266 s, 25368 tok |
| A | T2 |
3 | 0.542 (4.3/8) | 0.408 | 10.7 | 0.00 | 146 s, 21168 tok |
| A | T1T2 |
3 | 0.583 (4.7/8) | 0.424 | 11.0 | 0.00 | 236 s, 19328 tok |
| A | placeboA |
3 | 0.667 (5.3/8) | 0.229 | 24.3 | 0.00 | 356 s, 26850 tok |
| A | placeboB |
3 | 0.583 (4.7/8) | 0.245 | 19.0 | 0.00 | 272 s, 28238 tok |
| B | naive |
3 | 0.833 (6.7/8) | 0.366 | 18.3 | 1.67 | 119 s, 21710 tok |
| B | control |
3 | 0.625 (5.0/8) | 0.500 | 10.7 | 0.00 | 94 s, 12920 tok |
| B | T1 |
3 | 0.833 (6.7/8) | 0.645 | 10.3 | 0.00 | 232 s, 25122 tok |
| B | T2 |
3 | 0.708 (5.7/8) | 0.496 | 12.3 | 0.67 | 148 s, 20663 tok |
| B | T1T2 |
3 | 0.792 (6.3/8) | 0.552 | 12.3 | 0.00 | 271 s, 25826 tok |
| B | placeboA |
3 | 0.792 (6.3/8) | 0.349 | 20.0 | 0.33 | 299 s, 23286 tok |
| B | placeboB |
3 | 0.792 (6.3/8) | 0.465 | 14.0 | 0.00 | 246 s, 19363 tok |
| C | naive |
3 | 0.833 (6.7/8) | 0.278 | 24.7 | 2.00 | 99 s, 18196 tok |
| C | control |
3 | 0.792 (6.3/8) | 0.530 | 12.0 | 0.00 | 247 s, 15691 tok |
| C | T1 |
3 | 0.792 (6.3/8) | 0.392 | 19.3 | 0.67 | 219 s, 24786 tok |
| C | T2 |
3 | 0.792 (6.3/8) | 0.487 | 13.0 | 0.00 | 135 s, 14588 tok |
| C | T1T2 |
3 | 0.792 (6.3/8) | 0.365 | 22.0 | 0.00 | 196 s, 18869 tok |
| C | placeboA |
3 | 0.833 (6.7/8) | 0.246 | 27.3 | 0.00 | 300 s, 22411 tok |
| C | placeboB |
3 | 0.875 (7.0/8) | 0.441 | 16.0 | 0.00 | 264 s, 23347 tok |
| D | naive |
3 | 0.875 (7.0/8) | 0.502 | 14.7 | 0.33 | 58 s, 16853 tok |
| D | control |
3 | 0.792 (6.3/8) | 0.683 | 9.3 | 0.00 | 42 s, 21677 tok |
| D | T1 |
3 | 0.833 (6.7/8) | 0.468 | 14.7 | 0.00 | 115 s, 19761 tok |
| D | T2 |
3 | 0.875 (7.0/8) | 0.586 | 12.0 | 0.00 | 72 s, 12120 tok |
| D | T1T2 |
3 | 0.792 (6.3/8) | 0.513 | 12.3 | 0.00 | 168 s, 14800 tok |
| D | placeboA |
3 | 0.833 (6.7/8) | 0.420 | 16.3 | 0.00 | 146 s, 15094 tok |
| D | placeboB |
3 | 0.875 (7.0/8) | 0.492 | 14.3 | 0.00 | 85 s, 12916 tok |
| E | naive |
3 | 0.917 (7.3/8) | 0.440 | 16.7 | 0.33 | 99 s, 12864 tok |
| E | control |
3 | 0.958 (7.7/8) | 0.678 | 11.3 | 0.33 | 97 s, 13263 tok |
| E | T1 |
3 | 0.958 (7.7/8) | 0.639 | 13.0 | 0.33 | 222 s, 17967 tok |
| E | T2 |
3 | 0.958 (7.7/8) | 0.616 | 12.7 | 0.00 | 100 s, 13071 tok |
| E | T1T2 |
3 | 0.958 (7.7/8) | 0.530 | 15.0 | 0.00 | 260 s, 19582 tok |
| E | placeboA |
3 | 1.000 (8.0/8) | 0.271 | 30.7 | 1.33 | 289 s, 20287 tok |
| E | placeboB |
3 | 0.917 (7.3/8) | 0.435 | 17.0 | 0.00 | 230 s, 17918 tok |
| F | naive |
3 | 0.000 (0.0/1) | 0.000 | 15.7 | 2.00 | 224 s, 17585 tok |
| F | control |
3 | 0.000 (0.0/1) | 0.000 | 6.0 | 0.00 | 175 s, 15758 tok |
| F | T1 |
2 | 0.000 (0.0/1) | 0.000 | 9.0 | 0.00 | 277 s, 29519 tok |
| F | T2 |
3 | 0.333 (0.3/1) | 0.042 | 7.7 | 0.00 | 188 s, 22135 tok |
| F | T1T2 |
2 | 0.000 (0.0/1) | 0.000 | 10.5 | 0.00 | 238 s, 27830 tok |
| F | placeboA |
3 | 0.000 (0.0/1) | 0.000 | 13.3 | 0.00 | 410 s, 25812 tok |
| F | placeboB |
3 | 0.667 (0.7/1) | 0.067 | 10.3 | 0.00 | 318 s, 20747 tok |
| G | naive |
3 | 0.667 (0.7/1) | 0.065 | 10.0 | 0.33 | 44 s, 10776 tok |
| G | control |
3 | 0.667 (0.7/1) | 0.114 | 6.0 | 0.00 | 59 s, 11629 tok |
| G | T1 |
3 | 1.000 (1.0/1) | 0.145 | 7.0 | 0.00 | 133 s, 15085 tok |
| G | T2 |
3 | 1.000 (1.0/1) | 0.159 | 6.3 | 0.00 | 82 s, 12552 tok |
| G | T1T2 |
3 | 1.000 (1.0/1) | 0.132 | 7.7 | 0.67 | 166 s, 16498 tok |
| G | placeboA |
3 | 1.000 (1.0/1) | 0.065 | 16.3 | 0.00 | 235 s, 18268 tok |
| G | placeboB |
3 | 1.000 (1.0/1) | 0.089 | 11.3 | 0.00 | 133 s, 18167 tok |
claude-fable-5
| topic | arm | n | recall | yield | asserted | false alarms | cost |
|---|---|---|---|---|---|---|---|
| A | naive |
1 | 0.625 (5.0/8) | 0.263 | 19.0 | 5.00 | 137 s, 9120 tok |
| A | control |
1 | 0.625 (5.0/8) | 0.208 | 24.0 | 5.00 | 100 s, 6642 tok |
| A | T1 |
1 | 0.500 (4.0/8) | 0.333 | 12.0 | 0.00 | 195 s, 13473 tok |
| A | T2 |
1 | 0.500 (4.0/8) | 0.210 | 19.0 | 2.00 | 132 s, 8712 tok |
| A | T1T2 |
1 | 0.625 (5.0/8) | 0.556 | 9.0 | 0.00 | 188 s, 12901 tok |
| A | placeboA |
1 | 0.875 (7.0/8) | 0.636 | 11.0 | 0.00 | 186 s, 13071 tok |
| A | placeboB |
1 | 1.000 (8.0/8) | 0.286 | 28.0 | 5.00 | 211 s, 14680 tok |
| B | naive |
1 | 0.875 (7.0/8) | 0.438 | 16.0 | 3.00 | 128 s, 8396 tok |
| B | control |
1 | 0.750 (6.0/8) | 0.400 | 15.0 | 3.00 | 87 s, 5367 tok |
| B | T1 |
1 | 1.000 (8.0/8) | 0.444 | 18.0 | 3.00 | 157 s, 10722 tok |
| B | T2 |
1 | 0.875 (7.0/8) | 0.636 | 11.0 | 0.00 | 102 s, 6685 tok |
| B | T1T2 |
1 | 1.000 (8.0/8) | 0.444 | 18.0 | 3.00 | 216 s, 15325 tok |
| B | placeboA |
1 | 0.750 (6.0/8) | 0.207 | 29.0 | 6.00 | 139 s, 9924 tok |
| B | placeboB |
1 | 1.000 (8.0/8) | 0.381 | 21.0 | 3.00 | 128 s, 8731 tok |
| C | naive |
1 | 0.750 (6.0/8) | 0.214 | 28.0 | 7.00 | 104 s, 6500 tok |
| C | control |
1 | 0.750 (6.0/8) | 0.667 | 9.0 | 0.00 | 88 s, 5662 tok |
| C | T1 |
1 | 0.750 (6.0/8) | 0.231 | 26.0 | 8.00 | 206 s, 14429 tok |
| C | T2 |
1 | 0.750 (6.0/8) | 0.600 | 10.0 | 1.00 | 130 s, 8830 tok |
| C | T1T2 |
1 | 0.875 (7.0/8) | 0.292 | 24.0 | 6.00 | 145 s, 9926 tok |
| C | placeboA |
1 | 0.875 (7.0/8) | 0.350 | 20.0 | 4.00 | 198 s, 14313 tok |
| C | placeboB |
1 | 1.000 (8.0/8) | 0.235 | 34.0 | 5.00 | 162 s, 11402 tok |
| D | naive |
1 | 0.875 (7.0/8) | 0.467 | 15.0 | 0.00 | 79 s, 4758 tok |
| D | control |
1 | 0.875 (7.0/8) | 0.700 | 10.0 | 1.00 | 58 s, 3284 tok |
| D | T1 |
1 | 1.000 (8.0/8) | 0.471 | 17.0 | 0.00 | 115 s, 7642 tok |
| D | T2 |
1 | 0.625 (5.0/8) | 0.714 | 7.0 | 1.00 | 69 s, 4207 tok |
| D | T1T2 |
1 | 0.750 (6.0/8) | 0.667 | 9.0 | 0.00 | 87 s, 5524 tok |
| D | placeboA |
1 | 0.875 (7.0/8) | 0.500 | 14.0 | 0.00 | 154 s, 10551 tok |
| D | placeboB |
1 | 0.875 (7.0/8) | 0.583 | 12.0 | 2.00 | 86 s, 5654 tok |
| E | naive |
1 | 1.000 (8.0/8) | 0.276 | 29.0 | 11.00 | 94 s, 5748 tok |
| E | control |
1 | 0.875 (7.0/8) | 0.500 | 14.0 | 2.00 | 75 s, 4483 tok |
| E | T1 |
1 | 0.750 (6.0/8) | 0.429 | 14.0 | 2.00 | 124 s, 7952 tok |
| E | T2 |
1 | 0.875 (7.0/8) | 0.389 | 18.0 | 2.00 | 86 s, 5231 tok |
| E | T1T2 |
1 | 1.000 (8.0/8) | 0.381 | 21.0 | 3.00 | 127 s, 8327 tok |
| E | placeboA |
1 | 1.000 (8.0/8) | 0.421 | 19.0 | 3.00 | 126 s, 8222 tok |
| E | placeboB |
1 | 0.875 (7.0/8) | 0.538 | 13.0 | 1.00 | 95 s, 5999 tok |
| F | naive |
1 | 1.000 (1.0/1) | 0.067 | 15.0 | 4.00 | 139 s, 9604 tok |
| F | control |
1 | 1.000 (1.0/1) | 0.167 | 6.0 | 0.00 | 119 s, 8029 tok |
| F | T1 |
1 | 1.000 (1.0/1) | 0.091 | 11.0 | 3.00 | 216 s, 15263 tok |
| F | T2 |
1 | 1.000 (1.0/1) | 0.167 | 6.0 | 0.00 | 123 s, 8453 tok |
| F | T1T2 |
1 | 1.000 (1.0/1) | 0.200 | 5.0 | 1.00 | 175 s, 12092 tok |
| F | placeboA |
1 | 1.000 (1.0/1) | 0.056 | 18.0 | 1.00 | 208 s, 14892 tok |
| F | placeboB |
1 | 1.000 (1.0/1) | 0.071 | 14.0 | 0.00 | 210 s, 14459 tok |
| G | naive |
1 | 1.000 (1.0/1) | 0.062 | 16.0 | 3.00 | 94 s, 6171 tok |
| G | control |
1 | 1.000 (1.0/1) | 0.200 | 5.0 | 0.00 | 90 s, 6358 tok |
| G | T1 |
1 | 1.000 (1.0/1) | 0.250 | 4.0 | 0.00 | 112 s, 7423 tok |
| G | T2 |
1 | 1.000 (1.0/1) | 0.143 | 7.0 | 0.00 | 112 s, 7465 tok |
| G | T1T2 |
1 | 1.000 (1.0/1) | 0.143 | 7.0 | 0.00 | 92 s, 6168 tok |
| G | placeboA |
1 | 1.000 (1.0/1) | 0.100 | 10.0 | 2.00 | 137 s, 9754 tok |
| G | placeboB |
1 | 1.000 (1.0/1) | 0.143 | 7.0 | 2.00 | 117 s, 8140 tok |
gpt-5-6-pro
| topic | arm | n | recall | yield | asserted | false alarms | cost |
|---|---|---|---|---|---|---|---|
| A | naive |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| A | control |
1 | 0.500 (4.0/8) | 0.400 | 10.0 | 0.00 | 1230 s broker, 10172 chars |
| A | T1 |
1 | 0.500 (4.0/8) | 0.333 | 12.0 | 0.00 | 1230 s broker, 28755 chars |
| A | T2 |
1 | 0.625 (5.0/8) | 0.455 | 11.0 | 0.00 | 1776 s broker, 14880 chars |
| A | T1T2 |
1 | 0.500 (4.0/8) | 0.286 | 14.0 | 0.00 | 1308 s broker, 34649 chars |
| A | placeboA |
1 | 0.875 (7.0/8) | 0.226 | 31.0 | 0.00 | 2346 s broker, 24094 chars |
| A | placeboB |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| B | naive |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| B | control |
1 | 0.750 (6.0/8) | 0.600 | 10.0 | 1.00 | 1176 s broker, 7090 chars |
| B | T1 |
1 | 0.875 (7.0/8) | 0.200 | 35.0 | 4.00 | 2454 s broker, 20741 chars |
| B | T2 |
1 | 0.750 (6.0/8) | 0.250 | 24.0 | 2.00 | 1080 s broker, 10335 chars |
| B | T1T2 |
1 | 0.875 (7.0/8) | 0.500 | 14.0 | 1.00 | 3600 s broker, 29123 chars |
| B | placeboA |
1 | 0.875 (7.0/8) | 0.412 | 17.0 | 0.00 | 121980 s broker, 16636 chars |
| B | placeboB |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| C | naive |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| C | control |
1 | 0.750 (6.0/8) | 0.462 | 13.0 | 0.00 | 1884 s broker, 9938 chars |
| C | T1 |
1 | 0.875 (7.0/8) | 0.241 | 29.0 | 0.00 | 2454 s broker, 20466 chars |
| C | T2 |
1 | 0.875 (7.0/8) | 0.583 | 12.0 | 0.00 | 2448 s broker, 10710 chars |
| C | T1T2 |
1 | 0.750 (6.0/8) | 0.353 | 17.0 | 0.00 | 3024 s broker, 27942 chars |
| C | placeboA |
0 | not run | not run | not run | not run | not run |
| C | placeboB |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| D | naive |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| D | control |
1 | 0.875 (7.0/8) | 0.700 | 10.0 | 0.00 | 1086 s broker, 8784 chars |
| D | T1 |
1 | 0.875 (7.0/8) | 0.778 | 9.0 | 0.00 | 1164 s broker, 14935 chars |
| D | T2 |
1 | 0.875 (7.0/8) | 0.304 | 23.0 | 0.00 | 121974 s broker, 11321 chars |
| D | T1T2 |
1 | 0.875 (7.0/8) | 0.333 | 21.0 | 0.00 | 3600 s broker, 22813 chars |
| D | placeboA |
1 | 0.875 (7.0/8) | 0.333 | 21.0 | 0.00 | 121974 s broker, 17071 chars |
| D | placeboB |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| E | naive |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| E | control |
1 | 1.000 (8.0/8) | 0.667 | 12.0 | 0.00 | 2346 s broker, 11389 chars |
| E | T1 |
1 | 1.000 (8.0/8) | 0.615 | 13.0 | 0.00 | 2346 s broker, 18359 chars |
| E | T2 |
1 | 0.875 (7.0/8) | 0.636 | 11.0 | 0.00 | 2940 s broker, 15081 chars |
| E | T1T2 |
1 | 0.875 (7.0/8) | 0.700 | 10.0 | 0.00 | 1308 s broker, 20407 chars |
| E | placeboA |
1 | 1.000 (8.0/8) | 0.267 | 30.0 | 1.00 | 3558 s broker, 24925 chars |
| E | placeboB |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| F | naive |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| F | control |
1 | 0.000 (0.0/1) | 0.000 | 7.0 | 0.00 | 1206 s broker, 6441 chars |
| F | T1 |
1 | 0.000 (0.0/1) | 0.000 | 10.0 | 1.00 | 2358 s broker, 17991 chars |
| F | T2 |
1 | 0.000 (0.0/1) | 0.000 | 11.0 | 1.00 | 2358 s broker, 9939 chars |
| F | T1T2 |
1 | 0.000 (0.0/1) | 0.000 | 11.0 | 1.00 | 4080 s broker, 20708 chars |
| F | placeboA |
1 | 1.000 (1.0/1) | 0.062 | 16.0 | 2.00 | 2454 s broker, 16680 chars |
| F | placeboB |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| G | naive |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |
| G | control |
1 | 1.000 (1.0/1) | 0.167 | 6.0 | 0.00 | 1776 s broker, 6250 chars |
| G | T1 |
1 | 1.000 (1.0/1) | 0.143 | 7.0 | 0.00 | 2958 s broker, 18765 chars |
| G | T2 |
1 | 1.000 (1.0/1) | 0.143 | 7.0 | 0.00 | 2952 s broker, 8729 chars |
| G | T1T2 |
1 | 1.000 (1.0/1) | 0.125 | 8.0 | 0.00 | 4080 s broker, 16023 chars |
| G | placeboA |
1 | 1.000 (1.0/1) | 0.062 | 16.0 | 0.00 | 2448 s broker, 15667 chars |
| G | placeboB |
0 | not run | not run | not run | not run | arm outside the Pro plan (four core arms plus the matched placebo) |