Tips

Prompt Tips Tested on Claude, GPT and Gemini

I got tired of hunting for prompting tips that actually work, so I built a place to collect and test them.

18 tips tested 10 models 4,938 answers scored

17 tips, from 18 tested claims: the examples family reports two claims on one page.

How to read this

Each tip rests on a claim tested as a paired comparison on a fixed task set. Across 54 tested pairs the verdicts are 23 In free drift, 16 Stable, 14 Unobservable, 1 Past the horizon.

In free drift: could not be told from no effect; Stable: held; Unobservable: the test could not measure it; Past the horizon: measured worse.

Filter tips

Showing 17 of 17 tips.

  • What the marks mean

    • Holds
    • No measured effect
    • Could not measure
    • Helped a little, under our bar
    • Scored worse

Showing 17 of 17 tips.

Every verdict, by model

One row per tip, one column per model, on the instrument that measured it. The two instruments are two tables and are never read as one set: the same model on the API and in Claude Code is two measurements, not two opinions.

Tips · tested via API
TipClaude Haiku 4.5 · GPT-5 mini · Gemini 3.1 Flash Lite · GPT-5.4 mini · API twin for the Codex run ·
Getting clean JSON backHolds · +1.000Holds · +1.000Holds · +1.000Not run
Instructions before or afterCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Using tags and formattingCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Showing examplesNo measured effect · +0.157No measured effect · +0.186No measured effect · +0.157Not run
Better step-by-step answersHolds · +0.200Holds · +0.340No measured effect · +0.100Not run
Stop made-up answersHolds · +1.000Holds · +0.800Holds · +1.000Not run
Instructions in long promptsCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Offering the model moneyCould not measure · +44.2Could not measure · +179.6Could not measure · +56.5Not run
Giving the model a job titleNo measured effect · -0.013No measured effect · +0.133No measured effect · -0.033Not run
Adding pressureNo measured effect · +0.033No measured effect · +0.111No measured effect · -0.011Not run
Being politeNo measured effect · -0.100No measured effect · -0.056No measured effect · -0.128Not run
Asking for a rewriteHolds · +0.300No measured effect · +0.022No measured effect · -0.067Not run
Rules for the whole chatCould not measure · +0.000No measured effect · +0.000No measured effect · +0.100Not run
Splitting big tasksHolds · +0.233Holds · +0.247No measured effect · +0.100Not run
Prompt, then reviseCould not measure · +0.000No measured effect · +0.029No measured effect · +0.000Not run
Questions about recent eventsScored worse · -0.233Holds · +0.411Holds · +0.300Holds · +0.592
Getting the length rightHolds · +0.701Holds · +0.952Holds · +0.684Holds · +0.812
Tips · tested in Claude Code
TipClaude Fable 5 · Claude Opus 5 · Claude Haiku 4.5 · Claude Sonnet 5 · Claude Fable 5.1 · Claude Haiku 4.5 · multi-turn ·
Getting clean JSON backHolds · +1.000Holds · +1.000Holds · +1.000Holds · +1.000Holds · +1.000Not run
Instructions before or afterCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Using tags and formattingCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Showing examplesNo measured effect · +0.180No measured effect · +0.183No measured effect · +0.174No measured effect · +0.183No measured effect · +0.186Not run
Better step-by-step answersNo measured effect · +0.100No measured effect · +0.060Holds · +0.300No measured effect · +0.100No measured effect · +0.100Not run
Stop made-up answersHolds · +0.760Holds · +0.700Holds · +0.780Holds · +0.620Holds · +0.240Not run
Instructions in long promptsCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Offering the model moneyCould not measure · +62.0Could not measure · +283.5Could not measure · +16.6Could not measure · +80.0Could not measure · +105.1Not run
Giving the model a job titleNot runNot runNot runNot runNot runNot run
Adding pressureNot runNot runNot runNot runNot runNot run
Being politeNot runNot runNot runNot runNot runNot run
Asking for a rewriteNot runNot runNot runNot runNot runNot run
Rules for the whole chatNot runNot runNot runNot runNot runCould not measure · +0.000
Splitting big tasksNot runNot runNot runNot runNot runHolds · +0.313
Prompt, then reviseNot runNot runNot runNot runNot runNo measured effect · +0.029
Questions about recent eventsNo measured effect · +0.078No measured effect · +0.064No measured effect · +0.025No measured effect · -0.008No measured effect · -0.031Not run
Getting the length rightHolds · +0.235Holds · +0.823Holds · +0.650Holds · +0.774Holds · +0.494Not run
Tips · tested in Codex
TipGPT-5.4 mini · Codex run on GPT-5.4 mini · GPT-5.6 Luna · Codex run on GPT-5.6 Luna · GPT-5.6 Terra · Codex run on GPT-5.6 Terra · GPT-5.4 mini · Codex run, replaying two earlier tests ·
Getting clean JSON backHolds · +1.000Holds · +1.000Holds · +1.000Not run
Instructions before or afterCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Using tags and formattingCould not measure · +0.000No measured effect · -0.100Could not measure · +0.000Not run
Showing examplesHolds · +0.200No measured effect · +0.129No measured effect · +0.186Not run
Better step-by-step answersNo measured effect · +0.100No measured effect · +0.100No measured effect · +0.100Not run
Stop made-up answersHolds · +0.600Holds · +0.700Holds · +0.800Not run
Instructions in long promptsCould not measure · +0.000Could not measure · +0.000Could not measure · +0.000Not run
Offering the model moneyCould not measure · +102.9Could not measure · +83.5Could not measure · +82.1Not run
Giving the model a job titleNot runNot runNot runNot run
Adding pressureNot runNot runNot runNot run
Being politeNot runNot runNot runNot run
Asking for a rewriteNot runNot runNot runNot run
Rules for the whole chatNot runNot runNot runNot run
Splitting big tasksNot runNot runNot runNot run
Prompt, then reviseNot runNot runNot runNot run
Questions about recent eventsNot runNot runNo measured effect · +0.100Holds · +0.542
Getting the length rightHolds · +0.748Holds · +0.369Holds · +0.290Holds · +0.781

Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Fable 5 and Fable 5.1 are separate columns: they are two versions and this site does not merge them.

Ordering is derived. The lead is the one claim Stable on every model. The co-lead is the most-sampled null, computed by the rule in the methodology. Neither is chosen by hand.