Model under test

claude-haiku-4-5

Every claim tested against this model, with the verdict computed from the published records. Version string as the API returned it: claude-haiku-4-5-20251001.

Of the 18 tips tested on this model, 7 moved the score by more than the threshold with an interval that excludes zero. The other 11 are not a null result: they are the tips that could not be told apart from no effect, or could not be measured on their scale at all. The fact worth leading with here: the disclosure line scored worse (down 23 points).

Verdicts

One row per claim. Coverage is valid records over records attempted, across both arms, on every row. Rows marked task-paired were measured on the wave 2 instrument, where the interval rests on the task pair rather than on the record, so their intervals are built differently from the rows above them.
ClaimDeltaIntervalOrbit
An explicit JSON schema in the prompt yields a higher valid-JSON rate than an unstructured instruction to respond in JSON.+1.0001.000 to 1.000Stable100 of 100 records
Instructions placed after a document outperform instructions placed before it for extraction accuracy.+0.0000.000 to 0.000Unobservable100 of 100 records
XML tag delimiters improve instruction compliance over markdown headers on a multi-constraint task.+0.0000.000 to 0.000Unobservable100 of 100 records
Three worked examples improve format compliance over no examples, on a task where the format is genuinely underdetermined by the instruction.+0.1860.143 to 0.228task-pairedIn free drift100 of 100 records
One worked example improves format compliance over no examples, on the same task.+0.1570.129 to 0.185task-pairedIn free drift100 of 100 records
The phrase think step by step improves accuracy on multi-step word problems.+0.2000.010 to 0.390Stable100 of 100 records
Permitting the answer I do not know reduces fabricated answers on unanswerable questions.+1.0001.000 to 1.000Stable100 of 100 records
Instruction placement at the end of a long prompt beats placement in the middle for compliance.+0.0000.000 to 0.000Unobservable100 of 100 records
An offered tip increases output length. Length is the deterministic proxy this pilot can measure; the community claim is about quality, and the claim page will say so.+44.229.6 to 58.7Unobservable100 of 100 records
Role assignment improves response quality on domain questions.-0.013-0.265 to 0.238In free drift59 of 60 records
Emotional stakes framing improves response quality.+0.033-0.195 to 0.261In free drift60 of 60 records
Politeness markers change response quality.-0.100-0.322 to 0.122In free drift60 of 60 records
Asking the model to critique then revise its own answer improves final quality over a single pass.+0.3000.084 to 0.516Stable60 of 60 records
Standing rules hold better when placed in the system prompt than in the user message.+0.0000.000 to 0.000task-pairedUnobservable20 of 20 records
Breaking a complex task into separate sequential prompts beats one combined prompt.+0.2330.134 to 0.333task-pairedStable100 of 100 records
A generic refine pass after the answer beats one careful prompt.+0.0000.000 to 0.000task-pairedUnobservable20 of 20 records
Telling the model its knowledge may be out of date reduces confident fabrication on questions whose answers postdate its training.-0.233-0.416 to -0.050task-pairedPast the horizon110 of 110 records
Asking for an exact word count gets you that word count.+0.7010.476 to 0.926task-pairedStable100 of 100 records

How this model was called

Vendor
anthropic
Calls in the record set
1646 across 2 waves1040 in launch, on 2026-08-12; 606 in wave 2, 2026-08-15 to 2026-08-17.
Sampling
Sent and acceptedSent and accepted on every call, at temperature 0.
Reasoning suppression sent
no thinking parameter sent; claude-haiku-4-5 does not reason by default
Thinking tokens billed
Not reportedThis vendor reports no thinking token figure at all, so an accepted suppression setting is consistent with suppression but does not prove it.
  • Launch: 1040 calls, 2026-08-12, from harness/results/runs-b.jsonl. Recorded as claude-haiku-4-5-20251001.
  • Wave 2: 606 calls, 2026-08-15 to 2026-08-17, from harness/results/runs-w2.jsonl. The records carry the requested name back, with no dated snapshot behind it.

7 of 18 rows above were measured on the wave 2 instrument, on task-paired intervals. Why the two differ.