Model under test

Claude Haiku 4.5

Every result this site holds for Claude Haiku 4.5. We tested this model two ways, through the API and inside Claude Code, and never average them.

Version strings recorded, as returned: claude-haiku-4-5, claude-haiku-4-5-20251001. One answer names a dated build: claude-haiku-4-5-20251001. That is the vendor resolving the name we asked for to its own build of it, so it is the model we asked for, written down.

Of the 18 tips tested on this model via the API, 7 moved the score by more than the threshold, and the bracket around that move stays clear of zero. The other 11 are not a null result: they are the tips that could not be told apart from no effect, or could not be measured on their scale at all. The fact worth leading with here: the disclosure line scored worse (down 23 points).

What was measured, and where

We ran 78 separate tests on this model, listed by test method below and never added together. A test run through the API and inside Claude Code are two different measurements, so a single figure across them would be one no run produced. How the test methods were compared.

Prompting tips, tested via the API

18 tests, tested via API.

  • Holds 6 of 18
  • Could not measure 6 of 18
  • No measured effect 5 of 18
  • Scored worse 1 of 18

Prompting tips, replayed in Claude Code

12 tests, tested in Claude Code, Single-turn replay.

  • Could not measure 5 of 12
  • Holds 4 of 12
  • No measured effect 3 of 12

Prompting tips, multi-turn, in Claude Code

3 tests, tested in Claude Code, Multi-turn run.

  • No measured effect 1 of 3
  • Holds 1 of 3
  • Could not measure 1 of 3

Skills, tested via the API

12 tests, tested via API.

  • No measured effect 6 of 12
  • Holds 6 of 12

Skills, in Claude Code: Second run, every task twice

8 tests, tested in Claude Code, Second run, every task twice.

  • No measured effect 5 of 8
  • Holds 3 of 8

Skills, in Claude Code: Wider run, three ways

25 tests, tested in Claude Code, Wider run, three ways.

  • No measured effect 20 of 25
  • Holds 5 of 25

What these figures can be set beside

Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.

Measured on
tested via API; tested in Claude Code.
Not directly comparable
GPT-5.6 Luna; GPT-5.6 Terra were tested inside Codex, which this model did not run. No figure below is theirs and none of theirs is set against these.

This model is what connects through the API and inside Claude Code

Claude Haiku 4.5 was tested through the API and inside Claude Code, so its own readings are what connect them. They are listed job by job below and are never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.

Show the readings, job by job (20 jobs)
  • Accessibility audit Claude Haiku 4.5 scored 40.0% (26.3% to 53.7%) tested via API, First run, with and without the skill, and 45.5% (35.0% to 56.0%) tested in Claude Code, Second run, every task twice, and 46.1% (34.1% to 58.1%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
  • Code review Claude Haiku 4.5 scored 58.3% (47.8% to 68.7%) tested in Claude Code, Wider run, three ways. That is one test method only, so it bridges nothing on its own.
  • On-page audit Claude Haiku 4.5 scored 23.9% (1.8% to 46.1%) tested via API, First run, with and without the skill, and 23.1% (6.6% to 39.6%) tested in Claude Code, Second run, every task twice, and 23.1% (3.6% to 42.6%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
  • Skill authoring Claude Haiku 4.5 scored 61.8% (52.3% to 71.3%) tested via API, First run, with and without the skill, and 70.0% (63.5% to 76.5%) tested in Claude Code, Second run, every task twice, and 67.3% (59.7% to 74.9%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
  • Spec writing Claude Haiku 4.5 scored 99.5% (98.7% to 100.0%†) tested via API, First run, with and without the skill, and 98.6% (97.7% to 99.6%) tested in Claude Code, Second run, every task twice, and 99.2% (98.7% to 99.7%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
  • Getting clean JSON back Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 0.0% (0.0% to 0.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Instructions before or after Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Using tags and formatting Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Showing examples Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested in Claude Code. That is one test method only, so it bridges nothing on its own.
  • Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Better step-by-step answers Claude Haiku 4.5 scored 30.0% (17.2% to 42.8%) tested via API, Main run, tip by tip, and 20.0% (8.8% to 31.2%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Stop made-up answers Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 22.0% (10.4% to 33.6%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Instructions in long prompts Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Offering the model money Claude Haiku 4.5 scored 204 words (199 words to 209 words) tested via API, Main run, tip by tip, and 241 words (236 words to 247 words) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Questions about recent events Claude Haiku 4.5 scored 85.6% (73.2% to 98.0%) tested via API, Main run, tip by tip, and 96.7% (90.1% to 100.0%†) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Getting the length right Claude Haiku 4.5 scored 24.7% (1.2% to 48.2%) tested via API, Main run, tip by tip, and 27.5% (0.0% to 54.9%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
  • Rules for the whole chat Claude Haiku 4.5 scored 100.0% (no bracket published) tested in Claude Code. That is one test method only, so it bridges nothing on its own.
  • Splitting big tasks Claude Haiku 4.5 scored 68.7% (no bracket published) tested in Claude Code. That is one test method only, so it bridges nothing on its own.
  • Prompt, then revise Claude Haiku 4.5 scored 95.7% (no bracket published) tested in Claude Code. That is one test method only, so it bridges nothing on its own.

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

The grader, separately
a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.

What this model does on its own, one job at a time

We measured this model on 20 kinds of work, and the scores for every one of them are below.

Show the figures, one job at a time

One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.

Of the 5 models measured on this work, Claude Haiku 4.5 is the cheapest one whose score this test could not tell apart from the best score.

What one task costs with no help, and what it scores Showing examples

  • tested via API
  • tested in Claude Code

Across: what one task costs with no help, in US dollars at list price. Each gridline is ten times the one before it. Up: deterministic pass rate, 0 to 1.

  1. Gemini 3.1 Flash Lite Main run, tip by tip tested via API $0.00052 per task scored 80.0%, 75.4% to 84.6%
  2. GPT-5 mini Main run, tip by tip tested via API $0.00066 per task scored 78.3%, 73.0% to 83.6%
  3. Claude Haiku 4.5 Main run, tip by tip tested via API $0.00207 per task scored 81.4%, 77.2% to 85.7%
  4. Claude Haiku 4.5 tested in Claude Code $0.00311 per task scored 81.4%, 77.2% to 85.7%
  5. Claude Sonnet 5 tested in Claude Code $0.01713 per task scored 81.4%, 77.2% to 85.7%
  6. Claude Opus 5 tested in Claude Code $0.01882 per task scored 81.4%, 77.2% to 85.7%
  7. Claude Fable 5.1 tested in Claude Code $0.05013 per task scored 81.7%, 77.7% to 85.7%
  8. Claude Fable 5 tested in Claude Code $0.08170 per task scored 82.3%, 78.6% to 86.0%

Measured, and not on this chart

  • GPT-5.4 mini Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
  • GPT-5.6 Luna Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
  • GPT-5.6 Terra Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
Each mark is one model on one job: what a task cost it with no help, against the score it got. The bar through a mark shows how much that score could move if we ran it again. Shape and colour say which test method the model was reached through. The two test methods are measured on their own and are never ranked against each other. Compare marks of one colour with each other; a mark of one colour sitting above a mark of another is not a result. The tables are the source, and every figure drawn here is printed above. A model measured more than once appears once per test, named by test. Two tests are two measurements and are never averaged into one mark. Names sit beside their marks where there is room for them, moved down where two would have overlapped, and on a narrow screen they are in the list above instead. The list pairs every name with its figures at every width.

This is one of 20 sets of work this model was measured on. The same chart for every one of them sits on the comparison page, and the tables below carry this model’s figures for all of them.

Accessibility audit · seeded accessibility defects

tested via API, First run, with and without the skill

Accessibility audit, tested via API, First run, with and without the skill. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5 mini 69.9% 54.2% to 85.5% 20 records over 10 items, pooled from 2 entries 68.9% rampstackco-accessibility-audit (of 2 entries) $0.00035 billed through the API 515 in, 111 out not measured
2Gemini 3.1 Flash Lite 63.3% 46.0% to 80.7% 20 records over 10 items, pooled from 2 entries 72.8% rampstackco-accessibility-audit (of 2 entries) $0.00022 billed through the API 528 in, 59 out not measured
3Claude Haiku 4.5 40.0% 26.3% to 53.7% 20 records over 10 items, pooled from 2 entries 69.9% rampstackco-accessibility-audit (of 2 entries) $0.00088 billed through the API 600 in, 55 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 20 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Accessibility audit · seeded accessibility defects

tested in Claude Code, Second run, every task twice

Accessibility audit, tested in Claude Code, Second run, every task twice. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Opus 5 87.3% 78.3% to 96.4% 40 records over 10 items, pooled from 2 entries 88.0% affaan-m-accessibility (of 2 entries) $0.01384 run on subscription; shown at API list price for comparison 1,892 in, 175 out 0 a gate holding at zero, not a measured zero
2Claude Sonnet 5 82.3% 73.9% to 90.7% 40 records over 10 items, pooled from 2 entries 82.3% affaan-m-accessibility (of 2 entries) $0.01118 run on subscription; shown at API list price for comparison 2,544 in, 237 out 0 a gate holding at zero, not a measured zero
3Claude Haiku 4.5 45.5% 35.0% to 56.0% 40 records over 10 items, pooled from 2 entries 69.0% rampstackco-accessibility-audit (of 2 entries) $0.00206 run on subscription; shown at API list price for comparison 1,530 in, 106 out 0 a gate holding at zero, not a measured zero
  • Claude Fable 5: not published. 71 of 76 of the no-help calls in this job were answered by a different model (Claude Opus 5). We hold the figure back until those calls are bought again and answered by the model we asked for. The calls that went astray are not a random few. They land on the very runs this column averages, so a score built from the ones that came back correctly would describe which model the platform chose to send, not what this model can do. Why some figures are held back, under the rule that governs it.

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Accessibility audit · seeded accessibility defects

tested in Claude Code, Wider run, three ways

Accessibility audit, tested in Claude Code, Wider run, three ways. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 87.9% 76.3% to 99.5% 40 records over 10 items, pooled from 4 entries 91.5% alirezarezvani-a11y-audit (of 4 entries) not measured not measured in, 835 out 755
2Claude Opus 5 87.1% 78.0% to 96.2% 40 records over 10 items, pooled from 4 entries 91.5% affaan-m-accessibility (of 4 entries) not measured not measured in, 92 out 0 a gate holding at zero, not a measured zero
3Claude Sonnet 5 80.2% 71.4% to 88.9% 40 records over 10 items, pooled from 4 entries 80.7% alirezarezvani-a11y-audit (of 4 entries) not measured not measured in, 114 out 0 a gate holding at zero, not a measured zero
4Claude Haiku 4.5 46.1% 34.1% to 58.1% 40 records over 10 items, pooled from 4 entries 62.8% rampstackco-accessibility-audit (of 4 entries) not measured not measured in, 56 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

A cost per task needs both sides of the call. Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Code review · seeded code defects

tested in Claude Code, Wider run, three ways

Code review, tested in Claude Code, Wider run, three ways. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Opus 5 93.5% 84.7% to 100.0%† 80 records over 10 items, pooled from 8 entries 95.5% addyosmani-code-review-and-quality (of 8 entries) not measured not measured in, 113 out 0 a gate holding at zero, not a measured zero
2Claude Fable 5.1 89.0% 83.2% to 94.8% 80 records over 10 items, pooled from 8 entries 93.5% tied: jeffallan-code-reviewer, mattpocock-code-review, open-gsd-code-review (of 8 entries) not measured not measured in, 640 out 555
3Claude Sonnet 5 74.9% 62.4% to 87.4% 80 records over 10 items, pooled from 8 entries 83.3% rampstackco-code-review-web (of 8 entries) not measured not measured in, 74 out 0 a gate holding at zero, not a measured zero
4Claude Haiku 4.5 58.3% 47.8% to 68.7% 80 records over 10 items, pooled from 8 entries 56.7% open-gsd-code-review (of 8 entries) not measured not measured in, 41 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

A cost per task needs both sides of the call. Only 0 of the 80 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

Only 0 of the 80 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

On-page audit · seeded on-page issues

tested via API, First run, with and without the skill

On-page audit, tested via API, First run, with and without the skill. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Gemini 3.1 Flash Lite 48.8% 27.1% to 70.5% 20 records over 10 items, pooled from 2 entries 70.7% rampstackco-seo-onpage (of 2 entries) $0.00023 billed through the API 599 in, 53 out not measured
2GPT-5 mini 44.9% 26.7% to 63.1% 20 records over 10 items, pooled from 2 entries 49.2% rampstackco-seo-onpage (of 2 entries) $0.00029 billed through the API 566 in, 75 out not measured
3Claude Haiku 4.5 23.9% 1.8% to 46.1% 20 records over 10 items, pooled from 2 entries 47.4% rampstackco-seo-onpage (of 2 entries) $0.00085 billed through the API 667 in, 37 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 20 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

On-page audit · seeded on-page issues

tested in Claude Code, Second run, every task twice

On-page audit, tested in Claude Code, Second run, every task twice. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Opus 5 88.5% 77.6% to 99.5% 40 records over 10 items, pooled from 2 entries 83.3% rampstackco-seo-onpage (of 2 entries) not measured not measured in, 141 out 0 a gate holding at zero, not a measured zero
2Claude Fable 5 88.0% 79.2% to 96.9% 40 records over 10 items, pooled from 2 entries 94.0% affaan-m-seo (of 2 entries) not measured not measured in, 1,203 out 1,059
3Claude Sonnet 5 59.0% 46.0% to 72.0% 40 records over 10 items, pooled from 2 entries 64.7% affaan-m-seo (of 2 entries) not measured not measured in, 171 out 0 a gate holding at zero, not a measured zero
4Claude Haiku 4.5 23.1% 6.6% to 39.6% 40 records over 10 items, pooled from 2 entries 40.8% rampstackco-seo-onpage (of 2 entries) not measured not measured in, 83 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

A cost per task needs both sides of the call. Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

On-page audit · seeded on-page issues

tested in Claude Code, Wider run, three ways

On-page audit, tested in Claude Code, Wider run, three ways. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 97.2% 93.5% to 100.0%† 30 records over 10 items, pooled from 3 entries 98.3% rampstackco-seo-onpage (of 3 entries) not measured not measured in, 561 out 491
2Claude Opus 5 88.1% 76.9% to 99.2% 30 records over 10 items, pooled from 3 entries 83.3% alirezarezvani-seo-audit (of 3 entries) not measured not measured in, 72 out 0 a gate holding at zero, not a measured zero
3Claude Sonnet 5 53.0% 43.8% to 62.2% 30 records over 10 items, pooled from 3 entries 66.5% affaan-m-seo (of 3 entries) not measured not measured in, 84 out 0 a gate holding at zero, not a measured zero
4Claude Haiku 4.5 23.1% 3.6% to 42.6% 30 records over 10 items, pooled from 3 entries 41.1% rampstackco-seo-onpage (of 3 entries) not measured not measured in, 43 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

A cost per task needs both sides of the call. Only 0 of the 30 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

Only 0 of the 30 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Skill authoring · specification requirements

tested via API, First run, with and without the skill

Skill authoring, tested via API, First run, with and without the skill. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Gemini 3.1 Flash Lite 87.3% 80.1% to 94.4% 20 records over 10 items, pooled from 2 entries 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) $0.00088 billed through the API 59 in, 580 out not measured
2GPT-5 mini 73.2% 60.9% to 85.5% 20 records over 10 items, pooled from 2 entries 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) $0.00297 billed through the API 61 in, 1,478 out not measured
3Claude Haiku 4.5 61.8% 52.3% to 71.3% 20 records over 10 items, pooled from 2 entries 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) $0.00618 billed through the API 68 in, 1,222 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 20 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Skill authoring · specification requirements

tested in Claude Code, Second run, every task twice

Skill authoring, tested in Claude Code, Second run, every task twice. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 22 records over 10 items, pooled from 2 entries 100.0% anthropics-skill-creator (of 2 entries) $0.19551 run on subscription; shown at API list price for comparison 628 in, 3,785 out 131
1Claude Opus 5 100.0% 100.0% to 100.0% 30 records over 10 items, pooled from 2 entries 100.0% rampstackco-skill-creation-walkthrough (of 2 entries) $0.12485 run on subscription; shown at API list price for comparison 627 in, 4,869 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 40 records over 10 items, pooled from 2 entries 92.7% anthropics-skill-creator (of 2 entries) $0.05962 run on subscription; shown at API list price for comparison 1,278 in, 3,719 out 0 a gate holding at zero, not a measured zero
2Claude Haiku 4.5 70.0% 63.5% to 76.5% 40 records over 10 items, pooled from 2 entries 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) $0.00996 run on subscription; shown at API list price for comparison 467 in, 1,898 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Skill authoring · specification requirements

tested in Claude Code, Wider run, three ways

Skill authoring, tested in Claude Code, Wider run, three ways. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 100.0% 100.0% to 100.0% 40 records over 10 items, pooled from 4 entries 100.0% tied: openclaw-skill-creator, yusufkaraaslan-skill-builder (of 4 entries) not measured not measured in, 3,144 out 137
1Claude Opus 5 100.0% 100.0% to 100.0% 40 records over 10 items, pooled from 4 entries 100.0% tied: anthropics-skill-creator, openclaw-skill-creator, rampstackco-skill-creation-walkthrough, yusufkaraaslan-skill-builder (of 4 entries) not measured not measured in, 2,322 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 40 records over 10 items, pooled from 4 entries 100.0% tied: anthropics-skill-creator, openclaw-skill-creator, yusufkaraaslan-skill-builder (of 4 entries) not measured not measured in, 1,899 out 0 a gate holding at zero, not a measured zero
2Claude Haiku 4.5 67.3% 59.7% to 74.9% 40 records over 10 items, pooled from 4 entries 100.0% tied: anthropics-skill-creator, openclaw-skill-creator, rampstackco-skill-creation-walkthrough, yusufkaraaslan-skill-builder (of 4 entries) not measured not measured in, 961 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

A cost per task needs both sides of the call. Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Spec writing · required sections and acceptance criteria

tested via API, First run, with and without the skill

Spec writing, tested via API, First run, with and without the skill. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 99.5% 98.7% to 100.0%† 20 records over 10 items, pooled from 2 entries 99.1% rampstackco-pm-spec-writing (of 2 entries) $0.00332 billed through the API 118 in, 641 out not measured
2Gemini 3.1 Flash Lite 99.1% 97.3% to 100.0%† 20 records over 10 items, pooled from 2 entries 99.1% tied: affaan-m-product-capability, rampstackco-pm-spec-writing (of 2 entries) $0.00080 billed through the API 106 in, 513 out not measured
3GPT-5 mini 94.5% 92.0% to 97.1% 20 records over 10 items, pooled from 2 entries 90.0% affaan-m-product-capability (of 2 entries) $0.00248 billed through the API 107 in, 1,227 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

0 of the 40 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Spec writing · required sections and acceptance criteria

tested in Claude Code, Second run, every task twice

Spec writing, tested in Claude Code, Second run, every task twice. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 98.6% 97.7% to 99.6% 40 records over 10 items, pooled from 2 entries 97.7% rampstackco-pm-spec-writing (of 2 entries) not measured not measured in, 1,145 out 0 a gate holding at zero, not a measured zero
2Claude Fable 5 95.7% 94.6% to 96.7% 40 records over 10 items, pooled from 2 entries 92.7% tied: affaan-m-product-capability, rampstackco-pm-spec-writing (of 2 entries) not measured not measured in, 3,234 out 38
3Claude Sonnet 5 94.8% 93.0% to 96.5% 40 records over 10 items, pooled from 2 entries 95.0% rampstackco-pm-spec-writing (of 2 entries) not measured not measured in, 2,948 out 0 a gate holding at zero, not a measured zero
4Claude Opus 5 92.5% 91.3% to 93.7% 40 records over 10 items, pooled from 2 entries 94.5% rampstackco-pm-spec-writing (of 2 entries) not measured not measured in, 5,009 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

A cost per task needs both sides of the call. Only 26 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

A cost per task needs both sides of the call. Only 30 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

A cost per task needs both sides of the call. Only 32 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

A cost per task needs both sides of the call. Only 34 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

Only 26 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Only 30 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Only 32 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Only 34 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Spec writing · required sections and acceptance criteria

tested in Claude Code, Wider run, three ways

Spec writing, tested in Claude Code, Wider run, three ways. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 99.2% 98.7% to 99.7% 60 records over 10 items, pooled from 6 entries 99.1% github-prd (of 6 entries) not measured not measured in, 574 out 0 a gate holding at zero, not a measured zero
2Claude Fable 5.1 97.9% 96.8% to 99.0% 60 records over 10 items, pooled from 6 entries 100.0% rampstackco-pm-spec-writing (of 6 entries) not measured not measured in, 2,066 out 110
3Claude Sonnet 5 96.2% 94.9% to 97.6% 60 records over 10 items, pooled from 6 entries 97.3% github-prd (of 6 entries) not measured not measured in, 1,459 out 0 a gate holding at zero, not a measured zero
4Claude Opus 5 92.3% 91.1% to 93.5% 60 records over 10 items, pooled from 6 entries 94.5% phuryn-create-prd (of 6 entries) not measured not measured in, 2,362 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

A cost per task needs both sides of the call. Only 0 of the 60 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.

Only 0 of the 60 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting clean JSON back

tested via API, Main run, tip by tip

Getting clean JSON back, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema not measured 277 in, 415 out not measured
1Gemini 3.1 Flash Lite 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema not measured 219 in, 386 out not measured
1GPT-5 mini 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema not measured 235 in, 393 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting clean JSON back

tested in Claude Code

Getting clean JSON back, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.04321 run on subscription; shown at API list price for comparison 1,496 in, 565 out 46
1Claude Fable 5.1 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.04114 run on subscription; shown at API list price for comparison 1,506 in, 522 out 0
1Claude Haiku 4.5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.00316 run on subscription; shown at API list price for comparison 1,102 in, 412 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.02287 run on subscription; shown at API list price for comparison 1,496 in, 616 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.01740 run on subscription; shown at API list price for comparison 3,126 in, 535 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions before or after

tested via API, Main run, tip by tip

Instructions before or after, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after not measured 981 in, 41 out not measured
1Gemini 3.1 Flash Lite 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after not measured 986 in, 41 out not measured
1GPT-5 mini 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after not measured 900 in, 71 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions before or after

tested in Claude Code

Instructions before or after, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.02645 run on subscription; shown at API list price for comparison 2,465 in, 36 out 0
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.02928 run on subscription; shown at API list price for comparison 2,748 in, 36 out 0
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.00201 run on subscription; shown at API list price for comparison 1,806 in, 41 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.01323 run on subscription; shown at API list price for comparison 2,465 in, 36 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.01282 run on subscription; shown at API list price for comparison 4,095 in, 36 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Using tags and formatting

tested via API, Main run, tip by tip

Using tags and formatting, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters not measured 749 in, 759 out not measured
1Gemini 3.1 Flash Lite 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters not measured 686 in, 368 out not measured
1GPT-5 mini 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters not measured 708 in, 771 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Using tags and formatting

tested in Claude Code

Using tags and formatting, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 48 records 100.0% C03-xml-delimiters $0.11688 run on subscription; shown at API list price for comparison 2,370 in, 1,864 out 543
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.17308 run on subscription; shown at API list price for comparison 2,193 in, 3,023 out 1,351
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.00533 run on subscription; shown at API list price for comparison 1,574 in, 752 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.04752 run on subscription; shown at API list price for comparison 2,183 in, 1,464 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.02477 run on subscription; shown at API list price for comparison 3,813 in, 889 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Claude Code

Showing examples, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.03131 run on subscription; shown at API list price for comparison 1,719 in, 283 out 132
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.02474 run on subscription; shown at API list price for comparison 1,729 in, 149 out 0
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.00187 run on subscription; shown at API list price for comparison 1,281 in, 118 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.01238 run on subscription; shown at API list price for comparison 1,719 in, 151 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.01233 run on subscription; shown at API list price for comparison 3,349 in, 152 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested via API, Main run, tip by tip

Showing examples, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 81.4% 77.2% to 85.7% 10 items 97.1% C04b-zero-vs-one not measured 971 in, 221 out not measured
2Gemini 3.1 Flash Lite 80.0% 75.4% to 84.6% 10 items 95.7% C04b-zero-vs-one not measured 871 in, 199 out not measured
3GPT-5 mini 78.3% 73.0% to 83.6% 10 items 96.9% C04b-zero-vs-one not measured 875 in, 218 out not measured

0 of the 60 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Claude Code

Showing examples, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 82.3% 78.6% to 86.0% 10 items 99.7% C04b-zero-vs-one $0.08170 run on subscription; shown at API list price for comparison 2,621 in, 1,110 out 872
2Claude Fable 5.1 81.7% 77.7% to 85.7% 10 items 100.0% C04b-zero-vs-one $0.05013 run on subscription; shown at API list price for comparison 2,590 in, 485 out 252
3Claude Haiku 4.5 81.4% 77.2% to 85.7% 10 items 94.3% C04b-zero-vs-one $0.00311 run on subscription; shown at API list price for comparison 1,961 in, 230 out 0 a gate holding at zero, not a measured zero
3Claude Opus 5 81.4% 77.2% to 85.7% 10 items 98.0% C04b-zero-vs-one $0.01882 run on subscription; shown at API list price for comparison 2,578 in, 237 out 0 a gate holding at zero, not a measured zero
3Claude Sonnet 5 81.4% 77.2% to 85.7% 10 items 97.7% C04b-zero-vs-one $0.01713 run on subscription; shown at API list price for comparison 4,534 in, 235 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested via API, Main run, tip by tip

Showing examples, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 81.4% 77.2% to 85.7% 10 items 100.0% C04b-zero-vs-three not measured 971 in, 221 out not measured
2Gemini 3.1 Flash Lite 80.0% 75.4% to 84.6% 10 items 98.6% C04b-zero-vs-three not measured 871 in, 199 out not measured
3GPT-5 mini 79.4% 74.3% to 84.6% 10 items 97.7% C04b-zero-vs-three not measured 875 in, 219 out not measured

0 of the 60 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Claude Code

Showing examples, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 82.0% 78.2% to 85.8% 10 items 100.0% C04b-zero-vs-three $0.07734 run on subscription; shown at API list price for comparison 2,578 in, 1,031 out 798
2Claude Sonnet 5 81.7% 77.7% to 85.7% 10 items 100.0% C04b-zero-vs-three $0.01713 run on subscription; shown at API list price for comparison 4,534 in, 236 out 0 a gate holding at zero, not a measured zero
3Claude Fable 5.1 81.4% 77.2% to 85.7% 10 items 100.0% C04b-zero-vs-three $0.05010 run on subscription; shown at API list price for comparison 2,590 in, 484 out 252
3Claude Haiku 4.5 81.4% 77.2% to 85.7% 10 items 98.9% C04b-zero-vs-three $0.00310 run on subscription; shown at API list price for comparison 1,961 in, 228 out 0 a gate holding at zero, not a measured zero
3Claude Opus 5 81.4% 77.2% to 85.7% 10 items 99.7% C04b-zero-vs-three $0.01882 run on subscription; shown at API list price for comparison 2,578 in, 237 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Better step-by-step answers

tested via API, Main run, tip by tip

Better step-by-step answers, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Gemini 3.1 Flash Lite 40.0% 26.3% to 53.7% 50 records 50.0% C05-think-step-by-step not measured 220 in, 11 out not measured
2Claude Haiku 4.5 30.0% 17.2% to 42.8% 50 records 50.0% C05-think-step-by-step not measured 273 in, 25 out not measured
3GPT-5 mini 16.0% 5.7% to 26.3% 50 records 50.0% C05-think-step-by-step not measured 237 in, 53 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Better step-by-step answers

tested in Claude Code

Better step-by-step answers, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 40.0% 26.3% to 53.7% 50 records 50.0% C05-think-step-by-step $0.02052 run on subscription; shown at API list price for comparison 1,444 in, 122 out 107
1Claude Fable 5.1 40.0% 26.3% to 53.7% 50 records 50.0% C05-think-step-by-step $0.02103 run on subscription; shown at API list price for comparison 1,454 in, 130 out 0
1Claude Opus 5 40.0% 26.3% to 53.7% 50 records 46.0% C05-think-step-by-step $0.00807 run on subscription; shown at API list price for comparison 1,444 in, 34 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 40.0% 26.3% to 53.7% 50 records 50.0% C05-think-step-by-step $0.00945 run on subscription; shown at API list price for comparison 3,074 in, 15 out 0 a gate holding at zero, not a measured zero
2Claude Haiku 4.5 20.0% 8.8% to 31.2% 50 records 50.0% C05-think-step-by-step $0.00135 run on subscription; shown at API list price for comparison 1,098 in, 50 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Stop made-up answers

tested via API, Main run, tip by tip

Stop made-up answers, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5 mini 20.0% 8.8% to 31.2% 50 records 100.0% C06-permit-idk not measured 231 in, 165 out not measured
2Claude Haiku 4.5 0.0% 0.0% to 0.0% 50 records 100.0% C06-permit-idk not measured 256 in, 271 out not measured
2Gemini 3.1 Flash Lite 0.0% 0.0% to 0.0% 50 records 100.0% C06-permit-idk not measured 222 in, 87 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Stop made-up answers

tested in Claude Code

Stop made-up answers, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 76.0% 64.0% to 88.0% 50 records 100.0% C06-permit-idk $0.03175 run on subscription; shown at API list price for comparison 1,450 in, 345 out 0
2Claude Sonnet 5 38.0% 24.4% to 51.6% 50 records 100.0% C06-permit-idk $0.01625 run on subscription; shown at API list price for comparison 3,070 in, 469 out 0 a gate holding at zero, not a measured zero
3Claude Opus 5 30.0% 17.2% to 42.8% 50 records 100.0% C06-permit-idk $0.02363 run on subscription; shown at API list price for comparison 1,440 in, 657 out 0 a gate holding at zero, not a measured zero
4Claude Fable 5 24.0% 12.0% to 36.0% 50 records 100.0% C06-permit-idk $0.03501 run on subscription; shown at API list price for comparison 1,440 in, 412 out 44
5Claude Haiku 4.5 22.0% 10.4% to 33.6% 50 records 100.0% C06-permit-idk $0.00340 run on subscription; shown at API list price for comparison 1,081 in, 463 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions in long prompts

tested via API, Main run, tip by tip

Instructions in long prompts, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end not measured 1,386 in, 878 out not measured
1Gemini 3.1 Flash Lite 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end not measured 1,249 in, 398 out not measured
1GPT-5 mini 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end not measured 1,274 in, 845 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions in long prompts

tested in Claude Code

Instructions in long prompts, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 33 records 100.0% C07-instruction-at-end $0.23998 run on subscription; shown at API list price for comparison 4,267 in, 3,946 out 783
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.16619 run on subscription; shown at API list price for comparison 6,464 in, 2,031 out 362
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.00664 run on subscription; shown at API list price for comparison 2,211 in, 885 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.08287 run on subscription; shown at API list price for comparison 3,170 in, 2,681 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.03215 run on subscription; shown at API list price for comparison 4,800 in, 1,183 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Offering the model money

tested via API, Main run, tip by tip

Offering the model money, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelAnswer length, on its ownBest assisted scoreCost per taskTokens per taskThinking per task
1Gemini 3.1 Flash Lite 519 words 506 to 531 50 answers 575 words C08-tip-length not measured 59 in, 3,513 out not measured
2GPT-5 mini 355 words 334 to 375 50 answers 534 words C08-tip-length not measured 90 in, 2,337 out not measured
3Claude Haiku 4.5 204 words 199 to 209 50 answers 248 words C08-tip-length not measured 103 in, 1,619 out not measured

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Offering the model money

tested in Claude Code

Offering the model money, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelAnswer length, on its ownBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Opus 5 544 words 509 to 579 50 answers 828 words C08-tip-length $0.15452 run on subscription; shown at API list price for comparison 1,247 in, 5,931 out 0 a gate holding at zero, not a measured zero
2Claude Fable 5.1 403 words 385 to 421 50 answers 508 words C08-tip-length $0.23338 run on subscription; shown at API list price for comparison 1,257 in, 4,416 out 3
3Claude Sonnet 5 377 words 362 to 392 50 answers 457 words C08-tip-length $0.07162 run on subscription; shown at API list price for comparison 2,877 in, 4,199 out 0 a gate holding at zero, not a measured zero
4Claude Fable 5 325 words 313 to 336 50 answers 387 words C08-tip-length $0.20286 run on subscription; shown at API list price for comparison 1,247 in, 3,808 out 86
5Claude Haiku 4.5 241 words 236 to 247 50 answers 258 words C08-tip-length $0.01031 run on subscription; shown at API list price for comparison 928 in, 1,877 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Questions about recent events

tested via API, Main run, tip by tip

Questions about recent events, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 85.6% 73.2% to 98.0% 30 items 62.2% C16-cutoff-disclosure not measured 58 in, 162 out not measured
2GPT-5 mini 48.6% 31.8% to 65.4% 30 items 89.7% C16-cutoff-disclosure not measured 52 in, 131 out not measured
3Gemini 3.1 Flash Lite 3.3% 0.0% to 9.9%† 30 items 33.3% C16-cutoff-disclosure not measured 44 in, 90 out not measured

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

0 of the 108 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Questions about recent events

tested in Claude Code

Questions about recent events, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 96.7% 90.1% to 100.0%† 30 items 99.2% C16-cutoff-disclosure $0.00220 run on subscription; shown at API list price for comparison 474 in, 345 out 0 a gate holding at zero, not a measured zero
2Claude Sonnet 5 93.3% 84.3% to 100.0%† 30 items 92.5% C16-cutoff-disclosure $0.01136 run on subscription; shown at API list price for comparison 1,448 in, 467 out 0 a gate holding at zero, not a measured zero
3Claude Fable 5.1 92.2% 83.0% to 100.0%† 30 items 89.2% C16-cutoff-disclosure $0.04170 run on subscription; shown at API list price for comparison 638 in, 706 out 347
4Claude Fable 5 87.8% 77.5% to 98.0% 30 items 95.6% C16-cutoff-disclosure $0.04206 run on subscription; shown at API list price for comparison 633 in, 715 out 336
5Claude Opus 5 80.3% 67.4% to 93.1% 30 items 86.7% C16-cutoff-disclosure $0.01566 run on subscription; shown at API list price for comparison 633 in, 500 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting the length right

tested via API, Main run, tip by tip

Getting the length right, tested via API, Main run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Gemini 3.1 Flash Lite 31.0% 5.8% to 56.1% 10 items 99.4% C17-exact-length not measured 120 in, 3,398 out not measured
2Claude Haiku 4.5 24.7% 1.2% to 48.2% 10 items 94.8% C17-exact-length not measured 164 in, 2,238 out not measured
3GPT-5 mini 1.6% 0.0% to 4.6%† 10 items 96.8% C17-exact-length not measured 139 in, 5,014 out not measured

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

0 of the 70 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Nothing was billed on this run for the version being priced here.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting the length right

tested in Claude Code

Getting the length right, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 76.0% 53.7% to 98.4% 10 items 99.5% C17-exact-length $0.16668 run on subscription; shown at API list price for comparison 1,768 in, 2,980 out 437
2Claude Fable 5.1 50.2% 21.9% to 78.5% 10 items 99.5% C17-exact-length $0.22184 run on subscription; shown at API list price for comparison 1,786 in, 4,080 out 531
3Claude Haiku 4.5 27.5% 0.0% to 54.9% 10 items 92.5% C17-exact-length $0.01432 run on subscription; shown at API list price for comparison 1,319 in, 2,599 out 0 a gate holding at zero, not a measured zero
4Claude Sonnet 5 17.8% 0.0% to 41.0%† 10 items 95.2% C17-exact-length $0.09689 run on subscription; shown at API list price for comparison 4,050 in, 5,649 out 0 a gate holding at zero, not a measured zero
5Claude Opus 5 15.0% 0.0% to 35.1%† 10 items 97.3% C17-exact-length $0.18907 run on subscription; shown at API list price for comparison 1,768 in, 7,209 out 0 a gate holding at zero, not a measured zero

This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.

† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Rules for the whole chat

tested in Claude Code

Rules for the whole chat, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 100.0% interval not measured 10 items 100.0% C13-system-prompt-placement $0.01660 run on subscription; shown at API list price for comparison 13,384 in, 643 out not measured

0 of the 10 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Splitting big tasks

tested in Claude Code

Splitting big tasks, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 68.7% interval not measured 10 items 100.0% C14-prompt-chaining $0.00517 run on subscription; shown at API list price for comparison 3,902 in, 255 out not measured

0 of the 60 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Prompt, then revise

tested in Claude Code

Prompt, then revise, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 95.7% interval not measured 10 items 98.6% C15-iterative-refinement $0.00083 run on subscription; shown at API list price for comparison 557 in, 55 out not measured

0 of the 10 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Thinking was refused, and the refusal held on every call

This model was set up to run with thinking turned off, and any thinking at all on it was declared a failure that stops the run. Across every answer it gave, the count is zero throughout. That is the switch holding rather than a measured zero, and it is not comparable with a model that was allowed to think.

Answers reporting a nonzero count
0 of 1,330
Thinking tokens recorded
0

We ran 18 tests on this model through the API. Every one is below, with the difference we measured and how sure we are of it.

Show the full breakdown, tested via API

What we found, tested via API

One row per claim. Coverage is valid records over records attempted, across both arms, on every row. Rows marked task-paired were measured on the wave 2 instrument, where the interval rests on the task pair rather than on the record, so their intervals are built differently from the rows above them.
ClaimDeltaIntervalOrbit
An explicit JSON schema in the prompt yields a higher valid-JSON rate than an unstructured instruction to respond in JSON.+1.0001.000 to 1.000Stable100 of 100 records
Instructions placed after a document outperform instructions placed before it for extraction accuracy.+0.0000.000 to 0.000Unobservable100 of 100 records
XML tag delimiters improve instruction compliance over markdown headers on a multi-constraint task.+0.0000.000 to 0.000Unobservable100 of 100 records
Three worked examples improve format compliance over no examples, on a task where the format is genuinely underdetermined by the instruction.+0.1860.143 to 0.228task-pairedIn free drift100 of 100 records
One worked example improves format compliance over no examples, on the same task.+0.1570.129 to 0.185task-pairedIn free drift100 of 100 records
The phrase think step by step improves accuracy on multi-step word problems.+0.2000.010 to 0.390Stable100 of 100 records
Permitting the answer I do not know reduces fabricated answers on unanswerable questions.+1.0001.000 to 1.000Stable100 of 100 records
Instruction placement at the end of a long prompt beats placement in the middle for compliance.+0.0000.000 to 0.000Unobservable100 of 100 records
An offered tip increases output length. Length is the deterministic proxy this pilot can measure; the community claim is about quality, and the claim page will say so.+44.229.6 to 58.7Unobservable100 of 100 records
Role assignment improves response quality on domain questions.-0.013-0.265 to 0.238In free drift59 of 60 records
Emotional stakes framing improves response quality.+0.033-0.195 to 0.261In free drift60 of 60 records
Politeness markers change response quality.-0.100-0.322 to 0.122In free drift60 of 60 records
Asking the model to critique then revise its own answer improves final quality over a single pass.+0.3000.084 to 0.516Stable60 of 60 records
Standing rules hold better when placed in the system prompt than in the user message.+0.0000.000 to 0.000task-pairedUnobservable20 of 20 records
Breaking a complex task into separate sequential prompts beats one combined prompt.+0.2330.134 to 0.333task-pairedStable100 of 100 records
A generic refine pass after the answer beats one careful prompt.+0.0000.000 to 0.000task-pairedUnobservable20 of 20 records
Telling the model its knowledge may be out of date reduces confident fabrication on questions whose answers postdate its training.-0.233-0.416 to -0.050task-pairedPast the horizon110 of 110 post-cutoff records
Asking for an exact word count gets you that word count.+0.7010.476 to 0.926task-pairedStable100 of 100 records

How this model was called, tested via API

Vendor
anthropic
Calls behind these figures
1,646 across 2 waves 1,040 in launch, on 2026-08-12; 606 in wave 2, 2026-08-15 to 2026-08-17.
Sampling
Sent and accepted Sent and accepted on every call, at temperature 0.
Reasoning suppression sent
no thinking parameter sent; claude-haiku-4-5 does not reason by default
Thinking tokens billed
Not reported This vendor reports no thinking token figure at all, so an accepted suppression setting is consistent with suppression but does not prove it.
  • Launch: 1,040 calls, 2026-08-12, from this run’s committed answer file. Recorded as claude-haiku-4-5-20251001.
  • Wave 2: 606 calls, 2026-08-15 to 2026-08-17, from this run’s committed answer file. The answers carry the requested name back, with no dated snapshot behind it.

7 of 18 rows above come from the second wave, where the same task was run twice and the bracket is built from the two runs of each task rather than from every answer separately. Why the two differ.

This model was also run through Claude Code, on a path that bills nothing. How it was called, and what that means for every cost figure on this page, is below.

Show the full breakdown, tested in Claude Code

How this model was called, tested in Claude Code

On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this path returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.

These results are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.

Every figure on this page is worked out when the site is built, from the committed answers. Nothing here averages a figure from one test method with a figure from another, and the page carries no spread across them, because a model measured two ways has two results and not one.