Model under test
gpt-5-mini
Every claim tested against this model, with the verdict computed from the published records. Version string as the API returned it: gpt-5-mini-2025-08-07.
Of the 18 tips tested on this model, 6 moved the score by more than the threshold with an interval that excludes zero. The other 12 are not a null result: they are the tips that could not be told apart from no effect, or could not be measured on their scale at all. The fact worth leading with here: the strongest exact-length effect (up 95 points).
Verdicts
How this model was called
- Vendor
- openai
- Calls in the record set
- 1646 across 2 waves1040 in launch, on 2026-08-12; 606 in wave 2, 2026-08-15 to 2026-08-17.
- Sampling
- Rejected by the vendorRejected by the vendor. The fixed setting was sent, refused, and the call retried without it, so this model ran at its own default while the others ran at temperature 0.
- Reasoning suppression sent
- reasoning_effort=minimal
- Thinking tokens billed
- ReportedReported by the vendor. Nonzero on 0 of the 1040 launch calls; the wave 2 records do not carry the figure.
- Launch: 1040 calls, 2026-08-12, from harness/results/runs-b.jsonl. Recorded as gpt-5-mini-2025-08-07.
- Wave 2: 606 calls, 2026-08-15 to 2026-08-17, from harness/results/runs-w2.jsonl. The records carry the requested name back, with no dated snapshot behind it.
7 of 18 rows above were measured on the wave 2 instrument, on task-paired intervals. Why the two differ.
Where this model is documented by the party that ships it: OpenAI model documentation. Checked 2026-08-17. Get access: the OpenAI platform. Checked 2026-08-17.
Reading this model against another
Sampling is not held constant across vendors. This model ran at its own default because it refuses the fixed setting, so a comparison between this page and another model page compares two settings as well as two models. The caveat is permanent for the pilot record set.