Model under test

gemini-3.1-flash-lite

Every claim tested against this model, with the verdict computed from the published records. Version string as the API returned it: gemini-3.1-flash-lite.

Of the 18 tips tested on this model, 4 moved the score by more than the threshold with an interval that excludes zero. The other 14 are not a null result: they are the tips that could not be told apart from no effect, or could not be measured on their scale at all. The fact worth leading with here: fabricated 29 of 30 answers about recent events without a disclosure line.

Verdicts

One row per claim. Coverage is valid records over records attempted, across both arms, on every row. Rows marked task-paired were measured on the wave 2 instrument, where the interval rests on the task pair rather than on the record, so their intervals are built differently from the rows above them.
ClaimDeltaIntervalOrbit
An explicit JSON schema in the prompt yields a higher valid-JSON rate than an unstructured instruction to respond in JSON.+1.0001.000 to 1.000Stable100 of 100 records
Instructions placed after a document outperform instructions placed before it for extraction accuracy.+0.0000.000 to 0.000Unobservable100 of 100 records
XML tag delimiters improve instruction compliance over markdown headers on a multi-constraint task.+0.0000.000 to 0.000Unobservable100 of 100 records
Three worked examples improve format compliance over no examples, on a task where the format is genuinely underdetermined by the instruction.+0.1860.143 to 0.228task-pairedIn free drift100 of 100 records
One worked example improves format compliance over no examples, on the same task.+0.1570.129 to 0.185task-pairedIn free drift100 of 100 records
The phrase think step by step improves accuracy on multi-step word problems.+0.100-0.096 to 0.296In free drift100 of 100 records
Permitting the answer I do not know reduces fabricated answers on unanswerable questions.+1.0001.000 to 1.000Stable100 of 100 records
Instruction placement at the end of a long prompt beats placement in the middle for compliance.+0.0000.000 to 0.000Unobservable100 of 100 records
An offered tip increases output length. Length is the deterministic proxy this pilot can measure; the community claim is about quality, and the claim page will say so.+56.536.3 to 76.7Unobservable100 of 100 records
Role assignment improves response quality on domain questions.-0.033-0.200 to 0.133In free drift60 of 60 records
Emotional stakes framing improves response quality.-0.011-0.195 to 0.173In free drift60 of 60 records
Politeness markers change response quality.-0.128-0.316 to 0.060In free drift50 of 60 records
Asking the model to critique then revise its own answer improves final quality over a single pass.-0.067-0.236 to 0.103In free drift60 of 60 records
Standing rules hold better when placed in the system prompt than in the user message.+0.100-0.096 to 0.296task-pairedIn free drift20 of 20 records
Breaking a complex task into separate sequential prompts beats one combined prompt.+0.1000.000 to 0.200task-pairedIn free drift100 of 100 records
A generic refine pass after the answer beats one careful prompt.+0.0000.000 to 0.000task-pairedIn free drift20 of 20 records
Telling the model its knowledge may be out of date reduces confident fabrication on questions whose answers postdate its training.+0.3000.133 to 0.467task-pairedStable110 of 110 records
Asking for an exact word count gets you that word count.+0.6840.429 to 0.939task-pairedStable100 of 100 records

How this model was called

Vendor
gemini
Calls in the record set
1646 across 2 waves1040 in launch, on 2026-08-12; 606 in wave 2, 2026-08-15 to 2026-08-17.
Sampling
Sent and acceptedSent and accepted on every call, at temperature 0.
Reasoning suppression sent
thinkingConfig.thinkingBudget=0
Thinking tokens billed
Not reportedThis vendor reports no thinking token figure at all, so an accepted suppression setting is consistent with suppression but does not prove it.
  • Launch: 1040 calls, 2026-08-12, from harness/results/runs-b.jsonl. The records carry the requested name back, with no dated snapshot behind it.
  • Wave 2: 606 calls, 2026-08-15 to 2026-08-17, from harness/results/runs-w2.jsonl. The records carry the requested name back, with no dated snapshot behind it.

7 of 18 rows above were measured on the wave 2 instrument, on task-paired intervals. Why the two differ.