Model under test

gpt-5-mini

Every claim tested against this model, with the verdict computed from the published records. Version string as the API returned it: gpt-5-mini-2025-08-07.

Of the 18 tips tested on this model, 6 moved the score by more than the threshold with an interval that excludes zero. The other 12 are not a null result: they are the tips that could not be told apart from no effect, or could not be measured on their scale at all. The fact worth leading with here: the strongest exact-length effect (up 95 points).

Verdicts

One row per claim. Coverage is valid records over records attempted, across both arms, on every row. Rows marked task-paired were measured on the wave 2 instrument, where the interval rests on the task pair rather than on the record, so their intervals are built differently from the rows above them.
ClaimDeltaIntervalOrbit
An explicit JSON schema in the prompt yields a higher valid-JSON rate than an unstructured instruction to respond in JSON.+1.0001.000 to 1.000Stable100 of 100 records
Instructions placed after a document outperform instructions placed before it for extraction accuracy.+0.0000.000 to 0.000Unobservable100 of 100 records
XML tag delimiters improve instruction compliance over markdown headers on a multi-constraint task.+0.0000.000 to 0.000Unobservable100 of 100 records
Three worked examples improve format compliance over no examples, on a task where the format is genuinely underdetermined by the instruction.+0.1830.143 to 0.223task-pairedIn free drift100 of 100 records
One worked example improves format compliance over no examples, on the same task.+0.1860.140 to 0.232task-pairedIn free drift100 of 100 records
The phrase think step by step improves accuracy on multi-step word problems.+0.3400.166 to 0.514Stable100 of 100 records
Permitting the answer I do not know reduces fabricated answers on unanswerable questions.+0.8000.688 to 0.912Stable100 of 100 records
Instruction placement at the end of a long prompt beats placement in the middle for compliance.+0.0000.000 to 0.000Unobservable100 of 100 records
An offered tip increases output length. Length is the deterministic proxy this pilot can measure; the community claim is about quality, and the claim page will say so.+179.6144.5 to 214.7Unobservable100 of 100 records
Role assignment improves response quality on domain questions.+0.133-0.046 to 0.312In free drift60 of 60 records
Emotional stakes framing improves response quality.+0.111-0.056 to 0.278In free drift60 of 60 records
Politeness markers change response quality.-0.056-0.246 to 0.135In free drift60 of 60 records
Asking the model to critique then revise its own answer improves final quality over a single pass.+0.022-0.167 to 0.211In free drift60 of 60 records
Standing rules hold better when placed in the system prompt than in the user message.+0.0000.000 to 0.000task-pairedIn free drift20 of 20 records
Breaking a complex task into separate sequential prompts beats one combined prompt.+0.2470.162 to 0.332task-pairedStable100 of 100 records
A generic refine pass after the answer beats one careful prompt.+0.029-0.009 to 0.066task-pairedIn free drift20 of 20 records
Telling the model its knowledge may be out of date reduces confident fabrication on questions whose answers postdate its training.+0.4110.246 to 0.576task-pairedStable110 of 110 records
Asking for an exact word count gets you that word count.+0.9520.920 to 0.984task-pairedStable100 of 100 records

How this model was called

Vendor
openai
Calls in the record set
1646 across 2 waves1040 in launch, on 2026-08-12; 606 in wave 2, 2026-08-15 to 2026-08-17.
Sampling
Rejected by the vendorRejected by the vendor. The fixed setting was sent, refused, and the call retried without it, so this model ran at its own default while the others ran at temperature 0.
Reasoning suppression sent
reasoning_effort=minimal
Thinking tokens billed
ReportedReported by the vendor. Nonzero on 0 of the 1040 launch calls; the wave 2 records do not carry the figure.
  • Launch: 1040 calls, 2026-08-12, from harness/results/runs-b.jsonl. Recorded as gpt-5-mini-2025-08-07.
  • Wave 2: 606 calls, 2026-08-15 to 2026-08-17, from harness/results/runs-w2.jsonl. The records carry the requested name back, with no dated snapshot behind it.

7 of 18 rows above were measured on the wave 2 instrument, on task-paired intervals. Why the two differ.

Reading this model against another

Sampling is not held constant across vendors. This model ran at its own default because it refuses the fixed setting, so a comparison between this page and another model page compares two settings as well as two models. The caveat is permanent for the pilot record set.