Methodology
How a claim gets tested, and what the result can carry
The method is the authority here. It publishes in full, including the parts that limit what the results mean.
We call the tested statement a claim; the advice built on it is the tip.
By Andy Dunn.
Where claims come from
Claims are advice about prompting that circulates in practitioner communities: forums, documentation, and the working folklore of people who use these models daily. A claim qualifies for testing when it is stated widely enough to matter and specific enough to be wrong.
That a claim circulates is an input signal. It decides what gets tested. It is never evidence about whether the claim holds, and it never appears on a claim page as a number.
Making a claim testable
A claim becomes a test when it can be written as two prompts that differ in exactly one respect. One is the control arm, one is the treatment arm, and the difference between them is the whole of the manipulated variable. Everything else, including the task content, is identical by construction, because both arms are built by substituting the same task into two templates.
This is a paired comparison on a small fixed task set, not a controlled experiment on a representative sample. It supports no causal reading beyond the pairing.
The harness
18 claims ran against 3 models, both arms, on fixed task sets, producing 54 claim and model pairs from 4,938 records. 4 of the launch claims are rubric graded and the rest are scored by committed deterministic functions.
Those records arrived in two runs, on two instruments, and the table below is both of them. The split is set out under two instruments, one rule, because the difference changes how an interval on a wave 2 row was built and a reader comparing two rows should know which is which.
| Model | Version returned | Sampling | Reasoning sent |
|---|---|---|---|
| claude-haiku-4-5 | claude-haiku-4-5-20251001 | sent and accepted, temperature 0 | no thinking parameter sent; claude-haiku-4-5 does not reason by default |
| gpt-5-mini | gpt-5-mini-2025-08-07 | rejected by the vendor, ran at its default | reasoning_effort=minimal |
| gemini-3.1-flash-lite | gemini-3.1-flash-lite | sent and accepted, temperature 0 | thinkingConfig.thinkingBudget=0 |
A validity gate runs before scoring. A call that returned nothing usable, whether an empty output or a transport failure, is marked invalid, is never scored, and never enters a denominator. It is not a score of zero. Scoring an empty answer as zero was a real defect in an earlier instrument, and it is described below.
Two instruments, one rule
The claims were not all measured the same way. The launch run and the wave 2 run used different instruments, and rather than hide that behind one table, the difference is stated: it is a fact about how the numbers were made.
What did not change is the rule. Every verdict in the table, from either run, is assigned by the same frozen status_v1 classification, from the same file, with the same thresholds. Only the way an interval is built differs, and it differs because the two runs sampled differently.
| Run | Claims | Cells | Records | One observation is |
|---|---|---|---|---|
| Launch | 11 | 33 | 3,120 | one record |
| Wave 2 | 7 | 21 | 1,818 | one task pair |
Why a wave 2 interval is built differently
The launch run treats each record as one measurement, and its interval widens or narrows with the number of records behind it. That is right when every record is a separate look at the question.
Wave 2 is not that. Its task sets hold ten items, the design asked for fifty observations per cell, and the runner closed the gap by walking the same ten items five times. Every call ran at temperature 0, so the same prompt returned the same answer and scored the same: not five looks at a question, but one look written down five times. Counting those as fifty makes the interval about twice as narrow as the evidence supports, and the width of the interval is exactly what decides whether a verdict is allowed.
So on wave 2 the repeated passes collapse back to their item, the two arms are matched on the items they share, and the interval is built from the per-item differences. The unit is the pair, because the pair is what was independently measured. This is not a refinement made after the fact: it corrected a verdict that had already been published, which was then unpublished. The full account is in harness/PSEUDO-REPLICATION.md.
Coverage is the one figure that stays in records on both runs. It answers how much of what was attempted survived, which is a question about calls, not about observations, so the pairing is kept out of it.
Sets are calibrated before they are funded
A paid probe runs before a wave 2 task set is bought in full. It asks whether the control arm leaves the treatment arm anywhere to go: a set where the control already scores at the top cannot show an improvement, and one at the bottom cannot show a decline. Either way the run would buy a cell that measures nothing.
The control arm has to land between 0.1 and 0.9 of the scale to clear the gate. That is what the wave 2 probe checked, at 30 per set, on claude-haiku-4-5, on 2026-08-17, against a cap of 15 dollars. It cleared: C16-cutoff-disclosure came back at 86.0 percent of the ceiling, inside the walls, with room to move in either direction.
Saturation is the failure this is designed to catch, and the launch run has a worked example of missing it: the original three-example claim scored full marks in both arms on every model, which is not a null result but a task set that was too easy to separate anything. Its record is kept on the page that replaced it.
A claim's own target is not the threshold
Two numbers on this site look alike and are not. Each wave 2 claim carries a passCriterion, which is a design target: how large an effect the task set was built to be able to detect. It is set by the person building the set, and it answers whether the set is worth paying for.
The status_v1 threshold is what decides a verdict. It is fixed per scale, it is the same for every claim, and nothing in a claim file can change it. A claim can clear its own design target and still be In free drift, and that is not a contradiction: it means the set worked and the effect it found was real but smaller than the rule calls decisive.
Reading one as the other cost a re-run. The C14-prompt-chaining cell on gemini-3.1-flash-lite measured +0.100, which read against a design target of the same size looked like a boundary case, and it was authorised for more data. Under the rule it was never near a boundary, because the rule compares it against 0.2 and its verdict is In free drift. The extra records narrowed the interval exactly as more data should, and the verdict never moved. The account is in harness/PASS-CRITERION-VS-THRESHOLD.md.
One claim landed on both sides of the rule: telling a model its knowledge may be out of date cut confident fabrication on gpt-5-mini and gemini-3.1-flash-lite, and increased it on claude-haiku-4-5, which is the whole case for publishing every model separately rather than one verdict per claim.
The status_v1 rule
Every verdict on this site is assigned by one frozen rule. The site cannot vary it per claim, per model, or per page. A cell is one claim against one model, over both arms.
The comparison is the difference between the treatment mean and the control mean, with a 95 percent interval on that difference. A verdict requires the interval to exclude zero and the difference to clear a threshold.
| Scale | Pass threshold | Failure floor |
|---|---|---|
| Deterministic pass rate, 0 to 1 | +0.20 | -0.20 |
| Rubric mean, 1 to 5 | +0.30 | -0.30 |
| Word count, unbounded | none | none |
These thresholds are not chosen by taste. Each is the widest interval half-width the design produced on that scale, rounded up to a 0.05 step, over cells where at least one arm varied. A difference can only clear zero when it exceeds its own half-width, so the widest half-width observed is the smallest gap the design can be trusted to resolve. On the 0 to 1 scale most cells had no variance in either arm and could not contribute a half-width, so that threshold rests on a thin basis and is published as one.
Comparisons carry a fixed slack of 1e-9. Means are sums of floating point numbers, so a difference sitting exactly on a threshold can otherwise land either side of it in the sixteenth decimal place. The slack is far below the finest distinction the scorers can produce, so it cannot change a verdict the data supports.
The Orbit values
| Value | What it means | Cells |
|---|---|---|
| Stable | the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero | 16 |
| Past the horizon | the treatment arm scored below the control arm by more than the threshold, and the interval excludes zero | 1 |
| In free drift | no separation the design can resolve. Not evidence of no effect | 23 |
| Unobservable | the cell cannot be classified: an unbounded metric, or both arms at a bound of the scale | 14 |
| Decaying | a cell previously Stable that has fallen below the pass threshold on a later run | 0 |
Unobservable is tested before the thresholds. It covers two cases: an unbounded metric, where a fixed gap is not comparable to a gap on a bounded scorer and would be a number without a meaning; and both arms sitting at the same bound of the scale, where the task set did not separate the conditions. The second case is a fact about the instrument, not a measured null. Publishing it as In free drift would tell a reader the site measured no effect when it measured nothing at all.
A cell whose arms are both constant but at different bounds is not this case. Those cells are decisive, and they are classified normally.
Decaying requires the same cell measured on two runs. The record set is a single run, so no cell has a prior verdict to fall from and the value is unreachable today. It is defined so the rule does not have to change when re-testing begins, and the engine asserts in test that it is never emitted.
Coverage
Coverage is valid records over records attempted, for a cell, across both arms. It is computed separately from the verdict and never affects it. Where the two figures differ, both are shown, so a shortfall is visible rather than hidden behind a clean denominator.
How the lead claims are chosen
The lead is C06-permit-idk, the one claim that is Stable on every model in the set. That is a stated editorial constant, recorded so the choice is auditable rather than implicit.
The co-lead is computed, not chosen. It is the most-sampled resolved null: of every cell whose verdict is In free drift, keep those with complete coverage, rank by the number of valid records behind the null, and break ties on the widest interval. The intent is to surface the null the pilot looked hardest at and still could not resolve into an effect.
On the current records that selects C05-think-step-by-step on gemini-3.1-flash-lite, with 100 valid records behind a verdict of In free drift. It was selected from 10 clean nulls by that ranking, with no human input at any step.
Computing this rather than choosing it is the guarantee that the site does not cherry-pick its wins. The rule is designed to surface the sharpest available statement against the site's own thesis.
Limits of what this measured
Grading is the expense, not generation
The launch run's grader accounted for 64.6 percent of that run's spend, 4.4599 of 6.9071 dollars, across 710 launch verdicts; wave 2 ran deterministically with no grader calls. A rubric graded record cost about 19.9 times a deterministic one. This constrains how much rubric scored work the project can afford, and therefore how many claims can be tested that way.
The reasoning budget was a confound, and fixing it changed results
On an earlier instrument, with a tighter output budget, 11.6 percent of one vendor's calls spent the whole budget reasoning and returned empty text: 121 of 1040. Those were scored as zero, which looked like data, and the failure fell about 2.1 times as often on the treatment arm as the control arm, 82 against 39, because treatment prompts are longer. That correlated the manipulated variable with the failure and reversed the direction of every graded result for that vendor. With the budget raised, the empty-output residue is 0 across all 3 vendors. No figure from the earlier instrument is comparable to a figure published here.
There is no parity across vendors
One vendor rejects a temperature of 0 and ran at its own default on every call, with the rejection text recorded on each record. The others ran at 0. Reading one model column against another therefore compares two settings as well as two models. Separately, one vendor reports no thinking token figure at all, so an accepted suppression setting is consistent with suppression but does not prove it.
Model selection was a measurement decision
The pilot needed models that accept fixed sampling and do not reason by default, because a model that reasons before answering gives the control arm the behaviour the treatment arm asks for. Two current models from one vendor fail that test and were not used. The verdicts belong to the versions recorded on the records, not to any vendor's current best model.
Coverage is not uniform
One cell lost ten records to a transport failure across a 95 second window. One record carries a valid answer whose grader verdict could not be parsed, and is in neither the valid count nor the invalid count, so that cell's columns do not sum to its arm total. Both are named on the claim pages that carry them.
Saturation
11 of the 54 cells have both arms at a bound of the scale. Those task sets did not separate the conditions, and most of the claims that were deliberately hardened after an earlier run remained saturated afterwards. The hardening did not work. That is reported as a finding about the claims and the task design rather than smoothed into an inconclusive interval.
What is not published
The method publishes in full. The operational mechanics of how claims are captured, and the systems that run the pipeline, are not part of the method and do not publish.