Prompt Tips and AI Skills, Put to the Test

We test what people say works with AI, against a control, and publish it either way.

17 tips and 27 skills, tested on 8 models. Last run 2026-09-07.

Same score, less money

Best value, by job

On each job below, the cheapest model this test could not tell apart from the best one is drawn beside it. 5 models were measured on “One worked example”.

One worked example

Claude Fable 5

82.3%, about $0.0817 a task

Claude Haiku 4.5

81.4%, about $0.00311 a task

Three worked examples

Claude Fable 5

82.0%, about $0.0773 a task

Claude Haiku 4.5

81.4%, about $0.0031 a task

The ranges around these two scores overlap, so this test has not shown a difference between them. That is not the same as showing there is none: a real difference smaller than those ranges would look just like this. Both ranges are printed with the tables.

No reliable difference

What did not help

These were tested the same way as everything else. On every model that ran them, the run with the tip and the run without it came out too close to call.

Each one had room to score higher without the tip, so this is not a case of the task being too easy to show a difference.

The outline is the score without the tip and the filled bar is the score with it. What the bars pool, and what they do not.

Helped clearly

Tips that helped every model

The outline is the score without the tip and the filled bar is the score with it. What the bars pool, and what they do not.

Depends on the model

Skills

On Claude Haiku 4.5, 6 of the 25 skills tested helped clearly.

On Claude Fable 5, Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1, none did.

Tips held

Between two versions

12 of 12 held

Every one of them read the same both times.

Two skills read differently too: skill-creation-walkthrough and skill-creator.

Named, every time

Tested

8 models, three ways

Reached through the API, inside Claude Code and inside Codex. Results from the three are never added together.

Still running

Last run

357 calls

Finishing 2026-09-07.

Cost against score: “One worked example”

What one task costs with no help, and what it scores Showing examples

  • tested via API
  • tested in Claude Code

Across: what one task costs with no help, in US dollars at list price. Each gridline is ten times the one before it. Up: deterministic pass rate, 0 to 1. The up axis is fitted to what is drawn here rather than starting at the bottom of the scale, so its two ends are printed on it.

  1. Gemini 3.1 Flash Lite Main run, tip by tip tested via API $0.00052 per task scored 80.0%, 75.4% to 84.6%
  2. GPT-5 mini Main run, tip by tip tested via API $0.00066 per task scored 78.3%, 73.0% to 83.6%
  3. Claude Haiku 4.5 Main run, tip by tip tested via API $0.00207 per task scored 81.4%, 77.2% to 85.7%
  4. Claude Haiku 4.5 tested in Claude Code $0.00311 per task scored 81.4%, 77.2% to 85.7%
  5. Claude Sonnet 5 tested in Claude Code $0.01713 per task scored 81.4%, 77.2% to 85.7%
  6. Claude Opus 5 tested in Claude Code $0.01882 per task scored 81.4%, 77.2% to 85.7%
  7. Claude Fable 5.1 tested in Claude Code $0.05013 per task scored 81.7%, 77.7% to 85.7%
  8. Claude Fable 5 tested in Claude Code $0.08170 per task scored 82.3%, 78.6% to 86.0%
Each mark is one model on this job: what one task cost it with no help, against the score it got. The ringed mark is the cheaper of the two named above. Every job, with the tables behind it.

How it works

Every test runs the same task twice: once with the thing being tested, and once without it. The difference between those two runs is the whole of what is reported.

The models are named, and so is the version each one answered with. A result on one model is not published as a result on another.

When something published here turns out to be wrong, the correction goes up beside it and the original stays readable.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.