Model under test
claude-haiku-4-5
Every claim tested against this model, with the verdict computed from the published records. Version string as the API returned it: claude-haiku-4-5-20251001.
Of the 18 tips tested on this model, 7 moved the score by more than the threshold with an interval that excludes zero. The other 11 are not a null result: they are the tips that could not be told apart from no effect, or could not be measured on their scale at all. The fact worth leading with here: the disclosure line scored worse (down 23 points).
Verdicts
How this model was called
- Vendor
- anthropic
- Calls in the record set
- 1646 across 2 waves1040 in launch, on 2026-08-12; 606 in wave 2, 2026-08-15 to 2026-08-17.
- Sampling
- Sent and acceptedSent and accepted on every call, at temperature 0.
- Reasoning suppression sent
- no thinking parameter sent; claude-haiku-4-5 does not reason by default
- Thinking tokens billed
- Not reportedThis vendor reports no thinking token figure at all, so an accepted suppression setting is consistent with suppression but does not prove it.
- Launch: 1040 calls, 2026-08-12, from harness/results/runs-b.jsonl. Recorded as claude-haiku-4-5-20251001.
- Wave 2: 606 calls, 2026-08-15 to 2026-08-17, from harness/results/runs-w2.jsonl. The records carry the requested name back, with no dated snapshot behind it.
7 of 18 rows above were measured on the wave 2 instrument, on task-paired intervals. Why the two differ.
Where this model is documented by the party that ships it: Anthropic model documentation. Checked 2026-08-17. Get access: the Anthropic console. Checked 2026-08-17.