I got tired of hunting for prompting tips that actually work, so I built a place to collect and test them.
18 tips tested10 models4,938 answers scored
17 tips, from 18 tested claims: the examples family reports two claims on one page.
How to read this
Each tip rests on a claim tested as a paired comparison on a fixed task set. Across 54 tested pairs the verdicts are 23 In free drift, 16 Stable, 14 Unobservable, 1 Past the horizon.
How to use itAdd a line saying that if the text does not answer the question, replying I do not know is fine. Put it next to your question.
Pass rate up 80 to 100 points on all three models.
Models helped
Via API · Claude Haiku 4.5 · Holds · GPT-5 mini · Holds · Gemini 3.1 Flash Lite · Holds
In Claude Code · Claude Fable 5 · Holds · Claude Opus 5 · Holds · Claude Haiku 4.5 · Holds · Claude Sonnet 5 · Holds · Claude Fable 5.1 · Holds
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · Holds · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Holds · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Holds
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
How to use itAdd a line asking for the steps before the answer on any task with chained arithmetic or logic. Check that it helps on the model you use.
Pass rate up 20 points on Claude and 34 points on GPT. Inside the margin on Gemini.
Models helped
Via API · Claude Haiku 4.5 · Holds · GPT-5 mini · Holds · Gemini 3.1 Flash Lite · No measured effect
In Claude Code · Claude Fable 5 · No measured effect · Claude Opus 5 · No measured effect · Claude Haiku 4.5 · Holds · Claude Sonnet 5 · No measured effect · Claude Fable 5.1 · No measured effect
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · No measured effect · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · No measured effect · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · No measured effect
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
How to use itPut a small example of the structure in your prompt. List every field you need and what kind of value it holds.
Pass rate rose from 0% to 100% on all three models.
Models helped
Via API · Claude Haiku 4.5 · Holds · GPT-5 mini · Holds · Gemini 3.1 Flash Lite · Holds
In Claude Code · Claude Fable 5 · Holds · Claude Opus 5 · Holds · Claude Haiku 4.5 · Holds · Claude Sonnet 5 · Holds · Claude Fable 5.1 · Holds
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · Holds · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Holds · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Holds
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
How to use itPaste one example in the format you need. Add more only for formats that are tricky to describe in words. This helped a little on 9 of 10 models tested, under our bar for a clear win.
Pass rate up 18 to 19 points on all three models.
Models helped
Via API · Claude Haiku 4.5 · Helped a little, under our bar · GPT-5 mini · Helped a little, under our bar · Gemini 3.1 Flash Lite · Helped a little, under our bar · Claude Haiku 4.5 · Helped a little, under our bar · GPT-5 mini · Helped a little, under our bar · Gemini 3.1 Flash Lite · Helped a little, under our bar
In Claude Code · Claude Fable 5 · Could not measure · Claude Opus 5 · Could not measure · Claude Fable 5 · Helped a little, under our bar · Claude Opus 5 · Helped a little, under our bar · Claude Fable 5 · Helped a little, under our bar · Claude Opus 5 · Helped a little, under our bar · Claude Haiku 4.5 · Could not measure · Claude Sonnet 5 · Could not measure · Claude Haiku 4.5 · Helped a little, under our bar · Claude Sonnet 5 · Helped a little, under our bar · Claude Haiku 4.5 · Helped a little, under our bar · Claude Sonnet 5 · Helped a little, under our bar · Claude Fable 5.1 · Could not measure · Claude Fable 5.1 · Helped a little, under our bar · Claude Fable 5.1 · Helped a little, under our bar
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · Could not measure · GPT-5.4 mini · Codex run on GPT-5.4 mini · Holds · GPT-5.4 mini · Codex run on GPT-5.4 mini · Holds · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Could not measure · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Helped a little, under our bar · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Helped a little, under our bar · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Could not measure · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Helped a little, under our bar · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Helped a little, under our bar
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
How to use itAsk for the first step, then feed its answer into the next prompt. Best on tasks with several parts. Check that it helps on the model you use.
Pass rate up 10 to 25 points on all three models.
Models helped
Via API · Claude Haiku 4.5 · Holds · GPT-5 mini · Holds · Gemini 3.1 Flash Lite · No measured effect
In Claude Code · Claude Haiku 4.5 · multi-turn · Holds
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
How to use itSay exactly how many words you want rather than keep it short. Expect the answer to land within a few percent of the number you asked for.
Pass rate up 68 to 95 points on all three models.
Models helped
Via API · Claude Haiku 4.5 · Holds · GPT-5 mini · Holds · Gemini 3.1 Flash Lite · Holds · GPT-5.4 mini · API twin for the Codex run · Holds
In Claude Code · Claude Fable 5 · Holds · Claude Opus 5 · Holds · Claude Haiku 4.5 · Holds · Claude Sonnet 5 · Holds · Claude Fable 5.1 · Holds
In Codex · GPT-5.4 mini · Codex run, replaying two earlier tests · Holds · GPT-5.4 mini · Codex run on GPT-5.4 mini · Holds · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Holds · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Holds
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
+ Show the per-model numbers
Skip this
6 of 6Widely repeated, and showed no measured effect on any model tested.
Debunked
Giving the model a job title
You have probably heard you should tell the model to act as an expert.
What to do insteadOnly add it for questions about recent events, and check the marks: on Claude it made answers worse.
Pass rate up 41 points on GPT and 30 points on Gemini, and scored worse on Claude by 23 points.
Models helped
Via API · Claude Haiku 4.5 · Scored worse · GPT-5 mini · Holds · Gemini 3.1 Flash Lite · Holds · GPT-5.4 mini · API twin for the Codex run · Holds
In Claude Code · Claude Fable 5 · No measured effect · Claude Opus 5 · No measured effect · Claude Haiku 4.5 · No measured effect · Claude Sonnet 5 · No measured effect · Claude Fable 5.1 · No measured effect
In Codex · GPT-5.4 mini · Codex run, replaying two earlier tests · Holds · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · No measured effect
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
+ Show the per-model numbers
Inconclusive
4 of 4The instrument could not resolve these. No application claim is made.
Every model scored full marks with and without it, so the task set separated nothing.
Models helped
Via API · Claude Haiku 4.5 · Could not measure · GPT-5 mini · Could not measure · Gemini 3.1 Flash Lite · Could not measure
In Claude Code · Claude Fable 5 · Could not measure · Claude Opus 5 · Could not measure · Claude Haiku 4.5 · Could not measure · Claude Sonnet 5 · Could not measure · Claude Fable 5.1 · Could not measure
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · Could not measure · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Could not measure · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Could not measure
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
Every model scored full marks with and without it, so the task set separated nothing.
Models helped
Via API · Claude Haiku 4.5 · Could not measure · GPT-5 mini · Could not measure · Gemini 3.1 Flash Lite · Could not measure
In Claude Code · Claude Fable 5 · Could not measure · Claude Opus 5 · Could not measure · Claude Haiku 4.5 · Could not measure · Claude Sonnet 5 · Could not measure · Claude Fable 5.1 · Could not measure
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · Could not measure · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · No measured effect · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Could not measure
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
Every model scored full marks with and without it, so the task set separated nothing.
Models helped
Via API · Claude Haiku 4.5 · Could not measure · GPT-5 mini · Could not measure · Gemini 3.1 Flash Lite · Could not measure
In Claude Code · Claude Fable 5 · Could not measure · Claude Opus 5 · Could not measure · Claude Haiku 4.5 · Could not measure · Claude Sonnet 5 · Could not measure · Claude Fable 5.1 · Could not measure
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · Could not measure · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Could not measure · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Could not measure
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
Answer length rose by 44 to 180 words, but this metric has no threshold, so no verdict follows.
Models helped
Via API · Claude Haiku 4.5 · Could not measure · GPT-5 mini · Could not measure · Gemini 3.1 Flash Lite · Could not measure
In Claude Code · Claude Fable 5 · Could not measure · Claude Opus 5 · Could not measure · Claude Haiku 4.5 · Could not measure · Claude Sonnet 5 · Could not measure · Claude Fable 5.1 · Could not measure
In Codex · GPT-5.4 mini · Codex run on GPT-5.4 mini · Could not measure · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · Could not measure · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · Could not measure
One square is one model, on one test method. A named run is a second reading of the same model. Models we did not test on this tip have no square.
+ Show the per-model numbers
Showing 17 of 17 tips.
Every verdict, by model
One row per tip, one column per model, on the instrument that measured it. The two instruments are two tables and are never read as one set: the same model on the API and in Claude Code is two measurements, not two opinions.
Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Fable 5 and Fable 5.1 are separate columns: they are two versions and this site does not merge them.