Model under test
Claude Fable 5
Every result this site holds for Claude Fable 5. We tested it inside Claude Code.
Version strings recorded, as returned: claude-fable-5, claude-opus-5.
What was measured, and where
We ran 26 separate tests on this model, listed by test method below and never added together. How the test methods were compared.
Prompting tips, replayed in Claude Code
12 tests, tested in Claude Code, Single-turn replay.
- Could not measure 5 of 12
- No measured effect 4 of 12
- Holds 3 of 12
Skills, in Claude Code: First run, single attempt
6 tests, tested in Claude Code, First run, single attempt.
- No measured effect 5 of 6
- Could not measure 1 of 6
Skills, in Claude Code: Second run, every task twice
8 tests, tested in Claude Code, Second run, every task twice.
- No measured effect 7 of 8
- Could not measure 1 of 8
What these figures can be set beside
Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.
- Measured on
- tested in Claude Code.
- Not directly comparable
- Gemini 3.1 Flash Lite; GPT-5 mini; GPT-5.4 mini were tested through the API, which this model did not run. No figure below is theirs and none of theirs is set against these.
- Not directly comparable
- GPT-5.4 mini; GPT-5.6 Luna; GPT-5.6 Terra were tested inside Codex, which this model did not run. No figure below is theirs and none of theirs is set against these.
The model measured on both, job by job
Claude Haiku 4.5 was tested through the API and inside Claude Code. Its readings are the only thing connecting them, and they are stated job by job, never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.
Show the readings, job by job (15 jobs)
- On-page audit Claude Haiku 4.5 scored 23.9% (1.8% to 46.1%) tested via API, First run, with and without the skill, and 23.1% (6.6% to 39.6%) tested in Claude Code, Second run, every task twice, and 23.1% (3.6% to 42.6%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Skill authoring Claude Haiku 4.5 scored 61.8% (52.3% to 71.3%) tested via API, First run, with and without the skill, and 70.0% (63.5% to 76.5%) tested in Claude Code, Second run, every task twice, and 67.3% (59.7% to 74.9%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Spec writing Claude Haiku 4.5 scored 99.5% (98.7% to 100.0%†) tested via API, First run, with and without the skill, and 98.6% (97.7% to 99.6%) tested in Claude Code, Second run, every task twice, and 99.2% (98.7% to 99.7%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Getting clean JSON back Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 0.0% (0.0% to 0.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Instructions before or after Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Using tags and formatting Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Showing examples Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested in Claude Code. That is one test method only, so it bridges nothing on its own.
- Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Better step-by-step answers Claude Haiku 4.5 scored 30.0% (17.2% to 42.8%) tested via API, Main run, tip by tip, and 20.0% (8.8% to 31.2%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Stop made-up answers Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 22.0% (10.4% to 33.6%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Instructions in long prompts Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Offering the model money Claude Haiku 4.5 scored 204 words (199 words to 209 words) tested via API, Main run, tip by tip, and 241 words (236 words to 247 words) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Questions about recent events Claude Haiku 4.5 scored 85.6% (73.2% to 98.0%) tested via API, Main run, tip by tip, and 96.7% (90.1% to 100.0%†) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Getting the length right Claude Haiku 4.5 scored 24.7% (1.2% to 48.2%) tested via API, Main run, tip by tip, and 27.5% (0.0% to 54.9%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
- The grader, separately
- a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.
What this model does on its own, one job at a time
We measured this model on 15 kinds of work, and the scores for every one of them are below.
Show the figures, one job at a time
One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.
Claude Fable 5 scored higher than all 4 other models measured on this work.
What one task costs with no help, and what it scores Showing examples
- tested via API
- tested in Claude Code
Across: what one task costs with no help, in US dollars at list price. Each gridline is ten times the one before it. Up: deterministic pass rate, 0 to 1.
- Gemini 3.1 Flash Lite Main run, tip by tip tested via API $0.00052 per task scored 80.0%, 75.4% to 84.6%
- GPT-5 mini Main run, tip by tip tested via API $0.00066 per task scored 78.3%, 73.0% to 83.6%
- Claude Haiku 4.5 Main run, tip by tip tested via API $0.00207 per task scored 81.4%, 77.2% to 85.7%
- Claude Haiku 4.5 tested in Claude Code $0.00311 per task scored 81.4%, 77.2% to 85.7%
- Claude Sonnet 5 tested in Claude Code $0.01713 per task scored 81.4%, 77.2% to 85.7%
- Claude Opus 5 tested in Claude Code $0.01882 per task scored 81.4%, 77.2% to 85.7%
- Claude Fable 5.1 tested in Claude Code $0.05013 per task scored 81.7%, 77.7% to 85.7%
- Claude Fable 5 tested in Claude Code $0.08170 per task scored 82.3%, 78.6% to 86.0%
Measured, and not on this chart
- GPT-5.4 mini Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
- GPT-5.6 Luna Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
- GPT-5.6 Terra Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
This is one of 15 sets of work this model was measured on. The same chart for every one of them sits on the comparison page, and the tables below carry this model’s figures for all of them.
On-page audit · seeded on-page issues
tested in Claude Code, First run, single attempt
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 92.3% 85.3% to 99.4% 20 records over 10 items, pooled from 2 entries | 93.3% affaan-m-seo (of 2 entries) | $0.03843 run on subscription; shown at API list price for comparison | 1,068 in, 555 out | 483 |
| 2 | Claude Opus 5 | 87.9% 76.7% to 99.1% 20 records over 10 items, pooled from 2 entries | 83.3% rampstackco-seo-onpage (of 2 entries) | $0.00710 run on subscription; shown at API list price for comparison | 1,068 in, 71 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
On-page audit · seeded on-page issues
tested in Claude Code, Second run, every task twice
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 88.5% 77.6% to 99.5% 40 records over 10 items, pooled from 2 entries | 83.3% rampstackco-seo-onpage (of 2 entries) | not measured | not measured in, 141 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Fable 5 | 88.0% 79.2% to 96.9% 40 records over 10 items, pooled from 2 entries | 94.0% affaan-m-seo (of 2 entries) | not measured | not measured in, 1,203 out | 1,059 |
| 3 | Claude Sonnet 5 | 59.0% 46.0% to 72.0% 40 records over 10 items, pooled from 2 entries | 64.7% affaan-m-seo (of 2 entries) | not measured | not measured in, 171 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Haiku 4.5 | 23.1% 6.6% to 39.6% 40 records over 10 items, pooled from 2 entries | 40.8% rampstackco-seo-onpage (of 2 entries) | not measured | not measured in, 83 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Skill authoring · specification requirements
tested in Claude Code, First run, single attempt
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 20 records over 10 items, pooled from 2 entries | 100.0% anthropics-skill-creator (of 2 entries) | $0.10060 run on subscription; shown at API list price for comparison | 313 in, 1,949 out | 64 |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.06061 run on subscription; shown at API list price for comparison | 313 in, 2,362 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Skill authoring · specification requirements
tested in Claude Code, Second run, every task twice
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 22 records over 10 items, pooled from 2 entries | 100.0% anthropics-skill-creator (of 2 entries) | $0.19551 run on subscription; shown at API list price for comparison | 628 in, 3,785 out | 131 |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 30 records over 10 items, pooled from 2 entries | 100.0% rampstackco-skill-creation-walkthrough (of 2 entries) | $0.12485 run on subscription; shown at API list price for comparison | 627 in, 4,869 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 40 records over 10 items, pooled from 2 entries | 92.7% anthropics-skill-creator (of 2 entries) | $0.05962 run on subscription; shown at API list price for comparison | 1,278 in, 3,719 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Haiku 4.5 | 70.0% 63.5% to 76.5% 40 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00996 run on subscription; shown at API list price for comparison | 467 in, 1,898 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Spec writing · required sections and acceptance criteria
tested in Claude Code, First run, single attempt
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 93.2% 91.2% to 95.2% 20 records over 10 items, pooled from 2 entries | 93.6% affaan-m-product-capability (of 2 entries) | $0.08369 run on subscription; shown at API list price for comparison | 386 in, 1,597 out | 20 |
| 2 | Claude Opus 5 | 92.3% 90.9% to 93.6% 20 records over 10 items, pooled from 2 entries | 93.6% rampstackco-pm-spec-writing (of 2 entries) | $0.05944 run on subscription; shown at API list price for comparison | 386 in, 2,300 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Spec writing · required sections and acceptance criteria
tested in Claude Code, Second run, every task twice
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 98.6% 97.7% to 99.6% 40 records over 10 items, pooled from 2 entries | 97.7% rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 1,145 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Fable 5 | 95.7% 94.6% to 96.7% 40 records over 10 items, pooled from 2 entries | 92.7% tied: affaan-m-product-capability, rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 3,234 out | 38 |
| 3 | Claude Sonnet 5 | 94.8% 93.0% to 96.5% 40 records over 10 items, pooled from 2 entries | 95.0% rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 2,948 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Opus 5 | 92.5% 91.3% to 93.7% 40 records over 10 items, pooled from 2 entries | 94.5% rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 5,009 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 26 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
A cost per task needs both sides of the call. Only 30 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
A cost per task needs both sides of the call. Only 32 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
A cost per task needs both sides of the call. Only 34 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 26 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Only 30 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Only 32 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Only 34 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting clean JSON back
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.04321 run on subscription; shown at API list price for comparison | 1,496 in, 565 out | 46 |
| 1 | Claude Fable 5.1 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.04114 run on subscription; shown at API list price for comparison | 1,506 in, 522 out | 0 |
| 1 | Claude Haiku 4.5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.00316 run on subscription; shown at API list price for comparison | 1,102 in, 412 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.02287 run on subscription; shown at API list price for comparison | 1,496 in, 616 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.01740 run on subscription; shown at API list price for comparison | 3,126 in, 535 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions before or after
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.02645 run on subscription; shown at API list price for comparison | 2,465 in, 36 out | 0 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.02928 run on subscription; shown at API list price for comparison | 2,748 in, 36 out | 0 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.00201 run on subscription; shown at API list price for comparison | 1,806 in, 41 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.01323 run on subscription; shown at API list price for comparison | 2,465 in, 36 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.01282 run on subscription; shown at API list price for comparison | 4,095 in, 36 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Using tags and formatting
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 48 records | 100.0% C03-xml-delimiters | $0.11688 run on subscription; shown at API list price for comparison | 2,370 in, 1,864 out | 543 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.17308 run on subscription; shown at API list price for comparison | 2,193 in, 3,023 out | 1,351 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.00533 run on subscription; shown at API list price for comparison | 1,574 in, 752 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.04752 run on subscription; shown at API list price for comparison | 2,183 in, 1,464 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.02477 run on subscription; shown at API list price for comparison | 3,813 in, 889 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.03131 run on subscription; shown at API list price for comparison | 1,719 in, 283 out | 132 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.02474 run on subscription; shown at API list price for comparison | 1,729 in, 149 out | 0 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.00187 run on subscription; shown at API list price for comparison | 1,281 in, 118 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.01238 run on subscription; shown at API list price for comparison | 1,719 in, 151 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.01233 run on subscription; shown at API list price for comparison | 3,349 in, 152 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 82.3% 78.6% to 86.0% 10 items | 99.7% C04b-zero-vs-one | $0.08170 run on subscription; shown at API list price for comparison | 2,621 in, 1,110 out | 872 |
| 2 | Claude Fable 5.1 | 81.7% 77.7% to 85.7% 10 items | 100.0% C04b-zero-vs-one | $0.05013 run on subscription; shown at API list price for comparison | 2,590 in, 485 out | 252 |
| 3 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 94.3% C04b-zero-vs-one | $0.00311 run on subscription; shown at API list price for comparison | 1,961 in, 230 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Opus 5 | 81.4% 77.2% to 85.7% 10 items | 98.0% C04b-zero-vs-one | $0.01882 run on subscription; shown at API list price for comparison | 2,578 in, 237 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Sonnet 5 | 81.4% 77.2% to 85.7% 10 items | 97.7% C04b-zero-vs-one | $0.01713 run on subscription; shown at API list price for comparison | 4,534 in, 235 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 82.0% 78.2% to 85.8% 10 items | 100.0% C04b-zero-vs-three | $0.07734 run on subscription; shown at API list price for comparison | 2,578 in, 1,031 out | 798 |
| 2 | Claude Sonnet 5 | 81.7% 77.7% to 85.7% 10 items | 100.0% C04b-zero-vs-three | $0.01713 run on subscription; shown at API list price for comparison | 4,534 in, 236 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Fable 5.1 | 81.4% 77.2% to 85.7% 10 items | 100.0% C04b-zero-vs-three | $0.05010 run on subscription; shown at API list price for comparison | 2,590 in, 484 out | 252 |
| 3 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 98.9% C04b-zero-vs-three | $0.00310 run on subscription; shown at API list price for comparison | 1,961 in, 228 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Opus 5 | 81.4% 77.2% to 85.7% 10 items | 99.7% C04b-zero-vs-three | $0.01882 run on subscription; shown at API list price for comparison | 2,578 in, 237 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Better step-by-step answers
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | $0.02052 run on subscription; shown at API list price for comparison | 1,444 in, 122 out | 107 |
| 1 | Claude Fable 5.1 | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | $0.02103 run on subscription; shown at API list price for comparison | 1,454 in, 130 out | 0 |
| 1 | Claude Opus 5 | 40.0% 26.3% to 53.7% 50 records | 46.0% C05-think-step-by-step | $0.00807 run on subscription; shown at API list price for comparison | 1,444 in, 34 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | $0.00945 run on subscription; shown at API list price for comparison | 3,074 in, 15 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Haiku 4.5 | 20.0% 8.8% to 31.2% 50 records | 50.0% C05-think-step-by-step | $0.00135 run on subscription; shown at API list price for comparison | 1,098 in, 50 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Stop made-up answers
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 76.0% 64.0% to 88.0% 50 records | 100.0% C06-permit-idk | $0.03175 run on subscription; shown at API list price for comparison | 1,450 in, 345 out | 0 |
| 2 | Claude Sonnet 5 | 38.0% 24.4% to 51.6% 50 records | 100.0% C06-permit-idk | $0.01625 run on subscription; shown at API list price for comparison | 3,070 in, 469 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Opus 5 | 30.0% 17.2% to 42.8% 50 records | 100.0% C06-permit-idk | $0.02363 run on subscription; shown at API list price for comparison | 1,440 in, 657 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Fable 5 | 24.0% 12.0% to 36.0% 50 records | 100.0% C06-permit-idk | $0.03501 run on subscription; shown at API list price for comparison | 1,440 in, 412 out | 44 |
| 5 | Claude Haiku 4.5 | 22.0% 10.4% to 33.6% 50 records | 100.0% C06-permit-idk | $0.00340 run on subscription; shown at API list price for comparison | 1,081 in, 463 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions in long prompts
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 33 records | 100.0% C07-instruction-at-end | $0.23998 run on subscription; shown at API list price for comparison | 4,267 in, 3,946 out | 783 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.16619 run on subscription; shown at API list price for comparison | 6,464 in, 2,031 out | 362 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.00664 run on subscription; shown at API list price for comparison | 2,211 in, 885 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.08287 run on subscription; shown at API list price for comparison | 3,170 in, 2,681 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.03215 run on subscription; shown at API list price for comparison | 4,800 in, 1,183 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Offering the model money
tested in Claude Code
| Rank | Model | Answer length, on its own | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 544 words 509 to 579 50 answers | 828 words C08-tip-length | $0.15452 run on subscription; shown at API list price for comparison | 1,247 in, 5,931 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Fable 5.1 | 403 words 385 to 421 50 answers | 508 words C08-tip-length | $0.23338 run on subscription; shown at API list price for comparison | 1,257 in, 4,416 out | 3 |
| 3 | Claude Sonnet 5 | 377 words 362 to 392 50 answers | 457 words C08-tip-length | $0.07162 run on subscription; shown at API list price for comparison | 2,877 in, 4,199 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Fable 5 | 325 words 313 to 336 50 answers | 387 words C08-tip-length | $0.20286 run on subscription; shown at API list price for comparison | 1,247 in, 3,808 out | 86 |
| 5 | Claude Haiku 4.5 | 241 words 236 to 247 50 answers | 258 words C08-tip-length | $0.01031 run on subscription; shown at API list price for comparison | 928 in, 1,877 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Questions about recent events
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 96.7% 90.1% to 100.0%† 30 items | 99.2% C16-cutoff-disclosure | $0.00220 run on subscription; shown at API list price for comparison | 474 in, 345 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Sonnet 5 | 93.3% 84.3% to 100.0%† 30 items | 92.5% C16-cutoff-disclosure | $0.01136 run on subscription; shown at API list price for comparison | 1,448 in, 467 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Fable 5.1 | 92.2% 83.0% to 100.0%† 30 items | 89.2% C16-cutoff-disclosure | $0.04170 run on subscription; shown at API list price for comparison | 638 in, 706 out | 347 |
| 4 | Claude Fable 5 | 87.8% 77.5% to 98.0% 30 items | 95.6% C16-cutoff-disclosure | $0.04206 run on subscription; shown at API list price for comparison | 633 in, 715 out | 336 |
| 5 | Claude Opus 5 | 80.3% 67.4% to 93.1% 30 items | 86.7% C16-cutoff-disclosure | $0.01566 run on subscription; shown at API list price for comparison | 633 in, 500 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting the length right
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 76.0% 53.7% to 98.4% 10 items | 99.5% C17-exact-length | $0.16668 run on subscription; shown at API list price for comparison | 1,768 in, 2,980 out | 437 |
| 2 | Claude Fable 5.1 | 50.2% 21.9% to 78.5% 10 items | 99.5% C17-exact-length | $0.22184 run on subscription; shown at API list price for comparison | 1,786 in, 4,080 out | 531 |
| 3 | Claude Haiku 4.5 | 27.5% 0.0% to 54.9% 10 items | 92.5% C17-exact-length | $0.01432 run on subscription; shown at API list price for comparison | 1,319 in, 2,599 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Sonnet 5 | 17.8% 0.0% to 41.0%† 10 items | 95.2% C17-exact-length | $0.09689 run on subscription; shown at API list price for comparison | 4,050 in, 5,649 out | 0 a gate holding at zero, not a measured zero |
| 5 | Claude Opus 5 | 15.0% 0.0% to 35.1%† 10 items | 97.3% C17-exact-length | $0.18907 run on subscription; shown at API list price for comparison | 1,768 in, 7,209 out | 0 a gate holding at zero, not a measured zero |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
What changed between two versions
Higher on 4 jobs, lower on 1, no measurable change on 4, could not be measured on 6.
Whether a figure moved because the version changed is not something this design can say. Neither version was run twice and thinking is not deterministic on this transport, so a moved figure and a re-run of the same version are not distinguishable here. Every difference below is reported as measured and never as explained.
Each version against the one before it, per job
Offering the model money
scores higher
Claude Fable 5 scored 324.7. Claude Fable 5.1 scored 403.2. The change is +78.5. It could sit between +61.3 and +95.7. Both versions did 10 tasks here.
word count, unbounded This scale has no fixed ceiling. It is built from the values here, so a position is a share of the largest figure on it.
On-page audit
scores higher
Claude Fable 5 scored 0.923. Claude Fable 5.1 scored 0.983. The change is +0.060. It could sit between +0.011 and +0.109. Both versions did 10 tasks here.
Skill authoring
could not be measured
Claude Fable 5 scored 1.000. Claude Fable 5.1 scored 1.000. The change is +0.000. It could sit between +0.000 and +0.000. Both versions did 10 tasks here.
Spec writing
scores higher
Claude Fable 5 scored 0.932. Claude Fable 5.1 scored 0.977. The change is +0.045. It could sit between +0.022 and +0.068. Both versions did 10 tasks here.
Getting clean JSON back
could not be measured
Claude Fable 5 scored 0.000. Claude Fable 5.1 scored 0.000. The change is +0.000. It could sit between +0.000 and +0.000. Both versions did 10 tasks here.
Instructions before or after
could not be measured
Claude Fable 5 scored 1.000. Claude Fable 5.1 scored 1.000. The change is +0.000. It could sit between +0.000 and +0.000. Both versions did 10 tasks here.
Using tags and formatting
could not be measured
Claude Fable 5 scored 1.000. Claude Fable 5.1 scored 1.000. The change is +0.000. It could sit between +0.000 and +0.000. Both versions did 10 tasks here.
Showing examples
could not be measured
Claude Fable 5 scored 1.000. Claude Fable 5.1 scored 1.000. The change is +0.000. It could sit between +0.000 and +0.000. Both versions did 10 tasks here.
Showing examples
no measurable change
Claude Fable 5 scored 0.823. Claude Fable 5.1 scored 0.817. The change is -0.005. It could sit between -0.014 and +0.005. Both versions did 10 tasks here.
Showing examples
no measurable change
Claude Fable 5 scored 0.820. Claude Fable 5.1 scored 0.814. The change is -0.002. It could sit between -0.007 and +0.002. Both versions did 10 tasks here.
Better step-by-step answers
no measurable change
Claude Fable 5 scored 0.400. Claude Fable 5.1 scored 0.400. The change is +0.000. It could sit between +0.000 and +0.000. Both versions did 10 tasks here.
Stop made-up answers
scores higher
Claude Fable 5 scored 0.240. Claude Fable 5.1 scored 0.760. The change is +0.520. It could sit between +0.324 and +0.716. Both versions did 10 tasks here.
Instructions in long prompts
could not be measured
Claude Fable 5 scored 1.000. Claude Fable 5.1 scored 1.000. The change is +0.000. It could sit between +0.000 and +0.000. Both versions did 7 tasks here.
Questions about recent events
no measurable change
Claude Fable 5 scored 0.878. Claude Fable 5.1 scored 0.922. The change is +0.030. It could sit between -0.081 and +0.141. Both versions did 30 tasks here.
Getting the length right
scores lower
Claude Fable 5 scored 0.760. Claude Fable 5.1 scored 0.502. The change is -0.264. It could sit between -0.502 and -0.027. Both versions did 10 tasks here.
deterministic pass rate, 0 to 1 A bar wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.
On-page audit
On on-page audit, Claude Fable 5.1 finds 98% of seeded on-page issues unaided against 92% for Claude Fable 5; scores higher.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 92.3% 85.3% to 99.4% 10 items | $0.03843 run on subscription; shown at API list price for comparison | 483 | tested in Claude Code, Skill cohort (one pass) |
| Claude Fable 5.1 | 98.3% 95.1% to 100.0%† 10 items | $0.03907 run on subscription; shown at API list price for comparison | 499 | tested in Claude Code, Fable 5.1, one pass |
Claude Fable 5.1 scored +6.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +1.1 and +10.9 points, which is why we call it scores higher.
Skill authoring
On skill authoring, Claude Fable 5.1 finds 100% of specification requirements unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 10 items | $0.10060 run on subscription; shown at API list price for comparison | 64 | tested in Claude Code, Skill cohort (one pass) |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 10 items | $0.15512 run on subscription; shown at API list price for comparison | 145 | tested in Claude Code, Fable 5.1, one pass |
Claude Fable 5.1 scored +0.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +0.0 and +0.0 points, which is why we call it could not be measured.
Spec writing
On spec writing, Claude Fable 5.1 finds 98% of required sections and acceptance criteria unaided against 93% for Claude Fable 5; scores higher.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 93.2% 91.2% to 95.2% 10 items | $0.08369 run on subscription; shown at API list price for comparison | 20 | tested in Claude Code, Skill cohort (one pass) |
| Claude Fable 5.1 | 97.7% 96.2% to 99.2% 10 items | $0.11240 run on subscription; shown at API list price for comparison | 85 | tested in Claude Code, Fable 5.1, one pass |
Claude Fable 5.1 scored +4.5 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +2.2 and +6.8 points, which is why we call it scores higher.
Getting clean JSON back
On getting clean JSON back, Claude Fable 5.1 scores 0% unaided against 0% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 0.0% 0.0% to 0.0% 50 records | $0.04321 run on subscription; shown at API list price for comparison | 46 | tested in Claude Code |
| Claude Fable 5.1 | 0.0% 0.0% to 0.0% 50 records | $0.04114 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Claude Fable 5.1 scored +0.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +0.0 and +0.0 points, which is why we call it could not be measured.
Instructions before or after
On instructions before or after, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | $0.02645 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.02928 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Claude Fable 5.1 scored +0.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +0.0 and +0.0 points, which is why we call it could not be measured.
Using tags and formatting
On using tags and formatting, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 48 records | $0.11688 run on subscription; shown at API list price for comparison | 543 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.17308 run on subscription; shown at API list price for comparison | 1,351 | tested in Claude Code |
Claude Fable 5.1 scored +0.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +0.0 and +0.0 points, which is why we call it could not be measured.
Showing examples
On showing examples, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | $0.03131 run on subscription; shown at API list price for comparison | 132 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.02474 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Claude Fable 5.1 scored +0.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +0.0 and +0.0 points, which is why we call it could not be measured.
Showing examples
On showing examples, Claude Fable 5.1 scores 82% unaided against 82% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 82.3% 78.6% to 86.0% 10 items | $0.08170 run on subscription; shown at API list price for comparison | 872 | tested in Claude Code |
| Claude Fable 5.1 | 81.7% 77.7% to 85.7% 10 items | $0.05013 run on subscription; shown at API list price for comparison | 252 | tested in Claude Code |
Claude Fable 5.1 scored −0.5 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between −1.4 and +0.5 points, which is why we call it no measurable change.
Showing examples
On showing examples, Claude Fable 5.1 scores 81% unaided against 82% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 82.0% 78.2% to 85.8% 10 items | $0.07734 run on subscription; shown at API list price for comparison | 798 | tested in Claude Code |
| Claude Fable 5.1 | 81.4% 77.2% to 85.7% 10 items | $0.05010 run on subscription; shown at API list price for comparison | 252 | tested in Claude Code |
Claude Fable 5.1 scored −0.2 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between −0.7 and +0.2 points, which is why we call it no measurable change.
Better step-by-step answers
On better step-by-step answers, Claude Fable 5.1 scores 40% unaided against 40% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 40.0% 26.3% to 53.7% 50 records | $0.02052 run on subscription; shown at API list price for comparison | 107 | tested in Claude Code |
| Claude Fable 5.1 | 40.0% 26.3% to 53.7% 50 records | $0.02103 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Claude Fable 5.1 scored +0.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +0.0 and +0.0 points, which is why we call it no measurable change.
Stop made-up answers
On stop made-up answers, Claude Fable 5.1 scores 76% unaided against 24% for Claude Fable 5; scores higher.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 24.0% 12.0% to 36.0% 50 records | $0.03501 run on subscription; shown at API list price for comparison | 44 | tested in Claude Code |
| Claude Fable 5.1 | 76.0% 64.0% to 88.0% 50 records | $0.03175 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Claude Fable 5.1 scored +52.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +32.4 and +71.6 points, which is why we call it scores higher.
Instructions in long prompts
On instructions in long prompts, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 33 records | $0.23998 run on subscription; shown at API list price for comparison | 783 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.16619 run on subscription; shown at API list price for comparison | 362 | tested in Claude Code |
Claude Fable 5.1 scored +0.0 points against Claude Fable 5, on the same 7 items both answered. Run it again and the gap would most likely land between +0.0 and +0.0 points, which is why we call it could not be measured.
Offering the model money
On offering the model money, Claude Fable 5.1 scores 403 words unaided against 325 words for Claude Fable 5; scores higher.
| Version | Answer length, on its own | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 325 words 313 to 336 50 answers | $0.20286 run on subscription; shown at API list price for comparison | 86 | tested in Claude Code |
| Claude Fable 5.1 | 403 words 385 to 421 50 answers | $0.23338 run on subscription; shown at API list price for comparison | 3 | tested in Claude Code |
Claude Fable 5.1 scored +7852.0 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between +6133.2 and +9570.8 points, which is why we call it scores higher.
Questions about recent events
On questions about recent events, Claude Fable 5.1 scores 92% unaided against 88% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 87.8% 77.5% to 98.0% 30 items | $0.04206 run on subscription; shown at API list price for comparison | 336 | tested in Claude Code |
| Claude Fable 5.1 | 92.2% 83.0% to 100.0%† 30 items | $0.04170 run on subscription; shown at API list price for comparison | 347 | tested in Claude Code |
Claude Fable 5.1 scored +3.0 points against Claude Fable 5, on the same 30 items both answered. Run it again and the gap would most likely land between −8.1 and +14.1 points, which is why we call it no measurable change.
Getting the length right
On getting the length right, Claude Fable 5.1 scores 50% unaided against 76% for Claude Fable 5; scores lower.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 76.0% 53.7% to 98.4% 10 items | $0.16668 run on subscription; shown at API list price for comparison | 437 | tested in Claude Code |
| Claude Fable 5.1 | 50.2% 21.9% to 78.5% 10 items | $0.22184 run on subscription; shown at API list price for comparison | 531 | tested in Claude Code |
Claude Fable 5.1 scored −26.4 points against Claude Fable 5, on the same 10 items both answered. Run it again and the gap would most likely land between −50.2 and −2.7 points, which is why we call it scores lower.
This is the first point in a version series, not a rate. One step between two versions supports no curve and no guess at a third. Claude Fable 5 and Claude Fable 5.1 price identically, so a cost difference between them is a token difference.
Thinking was recorded rather than refused, by prior registration
Alone among the models tested this way, this one was set up in advance to have its thinking COUNTED rather than turned off, so thinking on it is not a failure. We wrote that down before the run precisely so that a result on this model could not later be explained away, or quietly discarded, on the grounds that it thought. The figures beside this are a measurement; the zeroes on the other models are a switch holding, and the two are not the same kind of number.
- Answers reporting a nonzero count
- 1,048 of 1,358
- Thinking tokens recorded
- 148,834
The plan served a different model on some calls, and those calls were refused
Some calls asking for this model came back answered by another one: same prompts, same flags, no error, and a plausible score. Nothing in the reply announced it except the name of the model that actually answered, which we now read back and keep on every answer. An answer from a different model tells us about that model, under conditions we chose for this one, so we throw the score away. We do not throw the answer away, and its tokens still count against what the run cost, because the call was made whoever answered it.
- Calls served by another model
- 48 of 1,358 on the tips replay, spanning 28 runs across 10 distinct combinations of tip, task and version Served instead by Claude Opus 5, on 2026-08-31.
- Skill results held back rather than thinned
- 2 Each is published as held back, with its reason and how many runs of the task survived, on the skill’s own page.
- Prevented from
- 2026-08-31 We now read back the name of the model that actually answered, and keep it on every answer. An answer from a model other than the one we asked for is thrown out when it arrives and again when it is loaded, under the check that stops it.
This model was also run through Claude Code, on a path that bills nothing. How it was called, and what that means for every cost figure on this page, is below.
Show the full breakdown, tested in Claude Code
How this model was called, tested in Claude Code
On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this path returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.
These results are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.