Model under test

Claude Fable 5.1

Every result this site holds for Claude Fable 5.1, reported per instrument. It was measured tested in Claude Code, and the two are never averaged together.

Version string recorded, as returned: claude-fable-5-1. The records carry the requested name back, with no dated snapshot behind it.

What was measured, and where

20 cells in total, and they are counted per instrument below rather than pooled. A cell measured tested via API and a cell measured tested in Claude Code are two measurements, so a single figure across them would be one neither run produced. How the two instruments were compared.

Prompting claims, replayed in Claude Code

12 cells, tested in Claude Code, single-turn replay.

  • Unobservable 5 of 12
  • In free drift 4 of 12
  • Stable 3 of 12

Skills, in Claude Code: Fable 5.1, one pass

8 cells, tested in Claude Code, fable 5.1, one pass.

  • In free drift 6 of 8
  • Past the horizon 1 of 8
  • Unobservable 1 of 8

What this model does unaided, per task set

One table per task set and instrument, ranked by unaided score. Nothing here is averaged across a task set, an instrument or a run. Every model’s tables sit together on the comparison page.

Accessibility audit · seeded accessibility defects

tested in Claude Code, Fable 5.1, one pass

Accessibility audit, tested in Claude Code, Fable 5.1, one pass. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 85.8% 72.6% to 99.1% 20 records over 10 items, pooled from 2 entries 93.5% rampstackco-accessibility-audit (of 2 entries) $0.05119 run on subscription; shown at API list price for comparison 948 in, 834 out 755

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

On-page audit · seeded on-page issues

tested in Claude Code, Fable 5.1, one pass

On-page audit, tested in Claude Code, Fable 5.1, one pass. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 98.3% 95.1% to 100.0%† 20 records over 10 items, pooled from 2 entries 95.0% tied: affaan-m-seo, rampstackco-seo-onpage (of 2 entries) $0.03907 run on subscription; shown at API list price for comparison 1,070 in, 568 out 499

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Skill authoring · specification requirements

tested in Claude Code, Fable 5.1, one pass

Skill authoring, tested in Claude Code, Fable 5.1, one pass. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 100.0% 100.0% to 100.0% 20 records over 10 items, pooled from 2 entries 100.0% rampstackco-skill-creation-walkthrough (of 2 entries) $0.15512 run on subscription; shown at API list price for comparison 315 in, 3,039 out 145

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Spec writing · required sections and acceptance criteria

tested in Claude Code, Fable 5.1, one pass

Spec writing, tested in Claude Code, Fable 5.1, one pass. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 97.7% 96.2% to 99.2% 20 records over 10 items, pooled from 2 entries 99.1% rampstackco-pm-spec-writing (of 2 entries) $0.11240 run on subscription; shown at API list price for comparison 388 in, 2,170 out 85

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting clean JSON back

tested in Claude Code

Getting clean JSON back, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.04321 run on subscription; shown at API list price for comparison 1,496 in, 565 out 46
1Claude Fable 5.1 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.04114 run on subscription; shown at API list price for comparison 1,506 in, 522 out 0
1Claude Haiku 4.5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.00316 run on subscription; shown at API list price for comparison 1,102 in, 412 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.02287 run on subscription; shown at API list price for comparison 1,496 in, 616 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 0.0% 0.0% to 0.0% 50 records 100.0% C01-json-schema $0.01740 run on subscription; shown at API list price for comparison 3,126 in, 535 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions before or after

tested in Claude Code

Instructions before or after, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.02645 run on subscription; shown at API list price for comparison 2,465 in, 36 out 0
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.02928 run on subscription; shown at API list price for comparison 2,748 in, 36 out 0
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.00201 run on subscription; shown at API list price for comparison 1,806 in, 41 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.01323 run on subscription; shown at API list price for comparison 2,465 in, 36 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C02-instruction-after $0.01282 run on subscription; shown at API list price for comparison 4,095 in, 36 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Using tags and formatting

tested in Claude Code

Using tags and formatting, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 48 records 100.0% C03-xml-delimiters $0.11688 run on subscription; shown at API list price for comparison 2,370 in, 1,864 out 543
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.17308 run on subscription; shown at API list price for comparison 2,193 in, 3,023 out 1,351
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.00533 run on subscription; shown at API list price for comparison 1,574 in, 752 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.04752 run on subscription; shown at API list price for comparison 2,183 in, 1,464 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C03-xml-delimiters $0.02477 run on subscription; shown at API list price for comparison 3,813 in, 889 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Claude Code

Showing examples, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.03131 run on subscription; shown at API list price for comparison 1,719 in, 283 out 132
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.02474 run on subscription; shown at API list price for comparison 1,729 in, 149 out 0
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.00187 run on subscription; shown at API list price for comparison 1,281 in, 118 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.01238 run on subscription; shown at API list price for comparison 1,719 in, 151 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C04-three-shot-format $0.01233 run on subscription; shown at API list price for comparison 3,349 in, 152 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Claude Code

Showing examples, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 82.3% 78.6% to 86.0% 10 items 99.7% C04b-zero-vs-one $0.08170 run on subscription; shown at API list price for comparison 2,621 in, 1,110 out 872
2Claude Fable 5.1 81.7% 77.7% to 85.7% 10 items 100.0% C04b-zero-vs-one $0.05013 run on subscription; shown at API list price for comparison 2,590 in, 485 out 252
3Claude Haiku 4.5 81.4% 77.2% to 85.7% 10 items 94.3% C04b-zero-vs-one $0.00311 run on subscription; shown at API list price for comparison 1,961 in, 230 out 0 a gate holding at zero, not a measured zero
3Claude Opus 5 81.4% 77.2% to 85.7% 10 items 98.0% C04b-zero-vs-one $0.01882 run on subscription; shown at API list price for comparison 2,578 in, 237 out 0 a gate holding at zero, not a measured zero
3Claude Sonnet 5 81.4% 77.2% to 85.7% 10 items 97.7% C04b-zero-vs-one $0.01713 run on subscription; shown at API list price for comparison 4,534 in, 235 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Claude Code

Showing examples, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 82.0% 78.2% to 85.8% 10 items 100.0% C04b-zero-vs-three $0.07734 run on subscription; shown at API list price for comparison 2,578 in, 1,031 out 798
2Claude Sonnet 5 81.7% 77.7% to 85.7% 10 items 100.0% C04b-zero-vs-three $0.01713 run on subscription; shown at API list price for comparison 4,534 in, 236 out 0 a gate holding at zero, not a measured zero
3Claude Fable 5.1 81.4% 77.2% to 85.7% 10 items 100.0% C04b-zero-vs-three $0.05010 run on subscription; shown at API list price for comparison 2,590 in, 484 out 252
3Claude Haiku 4.5 81.4% 77.2% to 85.7% 10 items 98.9% C04b-zero-vs-three $0.00310 run on subscription; shown at API list price for comparison 1,961 in, 228 out 0 a gate holding at zero, not a measured zero
3Claude Opus 5 81.4% 77.2% to 85.7% 10 items 99.7% C04b-zero-vs-three $0.01882 run on subscription; shown at API list price for comparison 2,578 in, 237 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Better step-by-step answers

tested in Claude Code

Better step-by-step answers, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 40.0% 26.3% to 53.7% 50 records 50.0% C05-think-step-by-step $0.02052 run on subscription; shown at API list price for comparison 1,444 in, 122 out 107
1Claude Fable 5.1 40.0% 26.3% to 53.7% 50 records 50.0% C05-think-step-by-step $0.02103 run on subscription; shown at API list price for comparison 1,454 in, 130 out 0
1Claude Opus 5 40.0% 26.3% to 53.7% 50 records 46.0% C05-think-step-by-step $0.00807 run on subscription; shown at API list price for comparison 1,444 in, 34 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 40.0% 26.3% to 53.7% 50 records 50.0% C05-think-step-by-step $0.00945 run on subscription; shown at API list price for comparison 3,074 in, 15 out 0 a gate holding at zero, not a measured zero
2Claude Haiku 4.5 20.0% 8.8% to 31.2% 50 records 50.0% C05-think-step-by-step $0.00135 run on subscription; shown at API list price for comparison 1,098 in, 50 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Stop made-up answers

tested in Claude Code

Stop made-up answers, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5.1 76.0% 64.0% to 88.0% 50 records 100.0% C06-permit-idk $0.03175 run on subscription; shown at API list price for comparison 1,450 in, 345 out 0
2Claude Sonnet 5 38.0% 24.4% to 51.6% 50 records 100.0% C06-permit-idk $0.01625 run on subscription; shown at API list price for comparison 3,070 in, 469 out 0 a gate holding at zero, not a measured zero
3Claude Opus 5 30.0% 17.2% to 42.8% 50 records 100.0% C06-permit-idk $0.02363 run on subscription; shown at API list price for comparison 1,440 in, 657 out 0 a gate holding at zero, not a measured zero
4Claude Fable 5 24.0% 12.0% to 36.0% 50 records 100.0% C06-permit-idk $0.03501 run on subscription; shown at API list price for comparison 1,440 in, 412 out 44
5Claude Haiku 4.5 22.0% 10.4% to 33.6% 50 records 100.0% C06-permit-idk $0.00340 run on subscription; shown at API list price for comparison 1,081 in, 463 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions in long prompts

tested in Claude Code

Instructions in long prompts, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 100.0% 100.0% to 100.0% 33 records 100.0% C07-instruction-at-end $0.23998 run on subscription; shown at API list price for comparison 4,267 in, 3,946 out 783
1Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.16619 run on subscription; shown at API list price for comparison 6,464 in, 2,031 out 362
1Claude Haiku 4.5 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.00664 run on subscription; shown at API list price for comparison 2,211 in, 885 out 0 a gate holding at zero, not a measured zero
1Claude Opus 5 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.08287 run on subscription; shown at API list price for comparison 3,170 in, 2,681 out 0 a gate holding at zero, not a measured zero
1Claude Sonnet 5 100.0% 100.0% to 100.0% 50 records 100.0% C07-instruction-at-end $0.03215 run on subscription; shown at API list price for comparison 4,800 in, 1,183 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Offering the model money

tested in Claude Code

Offering the model money, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Opus 5 54422.0% 50905.4% to 57938.6% 50 records 82768.0% C08-tip-length $0.15452 run on subscription; shown at API list price for comparison 1,247 in, 5,931 out 0 a gate holding at zero, not a measured zero
2Claude Fable 5.1 40320.0% 38542.7% to 42097.3% 50 records 50828.0% C08-tip-length $0.23338 run on subscription; shown at API list price for comparison 1,257 in, 4,416 out 3
3Claude Sonnet 5 37712.0% 36224.5% to 39199.5% 50 records 45708.0% C08-tip-length $0.07162 run on subscription; shown at API list price for comparison 2,877 in, 4,199 out 0 a gate holding at zero, not a measured zero
4Claude Fable 5 32468.0% 31347.7% to 33588.3% 50 records 38670.0% C08-tip-length $0.20286 run on subscription; shown at API list price for comparison 1,247 in, 3,808 out 86
5Claude Haiku 4.5 24138.0% 23624.8% to 24651.2% 50 records 25800.0% C08-tip-length $0.01031 run on subscription; shown at API list price for comparison 928 in, 1,877 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Questions about recent events

tested in Claude Code

Questions about recent events, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Haiku 4.5 96.7% 90.1% to 100.0%† 30 items 99.2% C16-cutoff-disclosure $0.00220 run on subscription; shown at API list price for comparison 474 in, 345 out 0 a gate holding at zero, not a measured zero
2Claude Sonnet 5 93.3% 84.3% to 100.0%† 30 items 92.5% C16-cutoff-disclosure $0.01136 run on subscription; shown at API list price for comparison 1,448 in, 467 out 0 a gate holding at zero, not a measured zero
3Claude Fable 5.1 92.2% 83.0% to 100.0%† 30 items 89.2% C16-cutoff-disclosure $0.04170 run on subscription; shown at API list price for comparison 638 in, 706 out 347
4Claude Fable 5 87.8% 77.5% to 98.0% 30 items 95.6% C16-cutoff-disclosure $0.04206 run on subscription; shown at API list price for comparison 633 in, 715 out 336
5Claude Opus 5 80.3% 67.4% to 93.1% 30 items 86.7% C16-cutoff-disclosure $0.01566 run on subscription; shown at API list price for comparison 633 in, 500 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting the length right

tested in Claude Code

Getting the length right, tested in Claude Code. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1Claude Fable 5 76.0% 53.7% to 98.4% 10 items 99.5% C17-exact-length $0.16668 run on subscription; shown at API list price for comparison 1,768 in, 2,980 out 437
2Claude Fable 5.1 50.2% 21.9% to 78.5% 10 items 99.5% C17-exact-length $0.22184 run on subscription; shown at API list price for comparison 1,786 in, 4,080 out 531
3Claude Haiku 4.5 27.5% 0.0% to 54.9% 10 items 92.5% C17-exact-length $0.01432 run on subscription; shown at API list price for comparison 1,319 in, 2,599 out 0 a gate holding at zero, not a measured zero
4Claude Sonnet 5 17.8% 0.0% to 41.0%† 10 items 95.2% C17-exact-length $0.09689 run on subscription; shown at API list price for comparison 4,050 in, 5,649 out 0 a gate holding at zero, not a measured zero
5Claude Opus 5 15.0% 0.0% to 35.1%† 10 items 97.3% C17-exact-length $0.18907 run on subscription; shown at API list price for comparison 1,768 in, 7,209 out 0 a gate holding at zero, not a measured zero

A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.

† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

What changed between two versions

Higher on 4 task sets, lower on 1, no measurable change on 4, could not be measured on 6.

Whether a figure moved because the version changed is not something this design can say. Neither version was run twice and thinking is not deterministic on this transport, so a moved figure and a re-run of the same version are not distinguishable here. Every difference below is reported as measured and never as explained.

On-page audit

On on-page audit, Claude Fable 5.1 finds 98% of seeded on-page issues unaided against 92% for Claude Fable 5; scores higher.

On-page audit: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 92.3% 85.3% to 99.4% 10 items $0.03843 run on subscription; shown at API list price for comparison 483tested in Claude Code, Skill cohort (one pass)
Claude Fable 5.1 98.3% 95.1% to 100.0%† 10 items $0.03907 run on subscription; shown at API list price for comparison 499tested in Claude Code, Fable 5.1, one pass

Paired difference: +6.0 points +1.1 to +10.9, over 10 committed items both versions answered. scores higher.

Skill authoring

On skill authoring, Claude Fable 5.1 finds 100% of specification requirements unaided against 100% for Claude Fable 5; could not be measured.

Skill authoring: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 100.0% 100.0% to 100.0% 10 items $0.10060 run on subscription; shown at API list price for comparison 64tested in Claude Code, Skill cohort (one pass)
Claude Fable 5.1 100.0% 100.0% to 100.0% 10 items $0.15512 run on subscription; shown at API list price for comparison 145tested in Claude Code, Fable 5.1, one pass

Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.

Spec writing

On spec writing, Claude Fable 5.1 finds 98% of required sections and acceptance criteria unaided against 93% for Claude Fable 5; scores higher.

Spec writing: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 93.2% 91.2% to 95.2% 10 items $0.08369 run on subscription; shown at API list price for comparison 20tested in Claude Code, Skill cohort (one pass)
Claude Fable 5.1 97.7% 96.2% to 99.2% 10 items $0.11240 run on subscription; shown at API list price for comparison 85tested in Claude Code, Fable 5.1, one pass

Paired difference: +4.5 points +2.2 to +6.8, over 10 committed items both versions answered. scores higher.

Getting clean JSON back

On getting clean JSON back, Claude Fable 5.1 scores 0% unaided against 0% for Claude Fable 5; could not be measured.

Getting clean JSON back: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 0.0% 0.0% to 0.0% 50 records $0.04321 run on subscription; shown at API list price for comparison 46tested in Claude Code
Claude Fable 5.1 0.0% 0.0% to 0.0% 50 records $0.04114 run on subscription; shown at API list price for comparison 0tested in Claude Code

Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.

Instructions before or after

On instructions before or after, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.

Instructions before or after: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 100.0% 100.0% to 100.0% 50 records $0.02645 run on subscription; shown at API list price for comparison 0tested in Claude Code
Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records $0.02928 run on subscription; shown at API list price for comparison 0tested in Claude Code

Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.

Using tags and formatting

On using tags and formatting, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.

Using tags and formatting: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 100.0% 100.0% to 100.0% 48 records $0.11688 run on subscription; shown at API list price for comparison 543tested in Claude Code
Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records $0.17308 run on subscription; shown at API list price for comparison 1,351tested in Claude Code

Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.

Showing examples

On showing examples, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.

Showing examples: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 100.0% 100.0% to 100.0% 50 records $0.03131 run on subscription; shown at API list price for comparison 132tested in Claude Code
Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records $0.02474 run on subscription; shown at API list price for comparison 0tested in Claude Code

Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.

Showing examples

On showing examples, Claude Fable 5.1 scores 82% unaided against 82% for Claude Fable 5; no measurable change.

Showing examples: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 82.3% 78.6% to 86.0% 10 items $0.08170 run on subscription; shown at API list price for comparison 872tested in Claude Code
Claude Fable 5.1 81.7% 77.7% to 85.7% 10 items $0.05013 run on subscription; shown at API list price for comparison 252tested in Claude Code

Paired difference: −0.5 points −1.4 to +0.5, over 10 committed items both versions answered. no measurable change.

Showing examples

On showing examples, Claude Fable 5.1 scores 81% unaided against 82% for Claude Fable 5; no measurable change.

Showing examples: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 82.0% 78.2% to 85.8% 10 items $0.07734 run on subscription; shown at API list price for comparison 798tested in Claude Code
Claude Fable 5.1 81.4% 77.2% to 85.7% 10 items $0.05010 run on subscription; shown at API list price for comparison 252tested in Claude Code

Paired difference: −0.2 points −0.7 to +0.2, over 10 committed items both versions answered. no measurable change.

Better step-by-step answers

On better step-by-step answers, Claude Fable 5.1 scores 40% unaided against 40% for Claude Fable 5; no measurable change.

Better step-by-step answers: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 40.0% 26.3% to 53.7% 50 records $0.02052 run on subscription; shown at API list price for comparison 107tested in Claude Code
Claude Fable 5.1 40.0% 26.3% to 53.7% 50 records $0.02103 run on subscription; shown at API list price for comparison 0tested in Claude Code

Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. no measurable change.

Stop made-up answers

On stop made-up answers, Claude Fable 5.1 scores 76% unaided against 24% for Claude Fable 5; scores higher.

Stop made-up answers: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 24.0% 12.0% to 36.0% 50 records $0.03501 run on subscription; shown at API list price for comparison 44tested in Claude Code
Claude Fable 5.1 76.0% 64.0% to 88.0% 50 records $0.03175 run on subscription; shown at API list price for comparison 0tested in Claude Code

Paired difference: +52.0 points +32.4 to +71.6, over 10 committed items both versions answered. scores higher.

Instructions in long prompts

On instructions in long prompts, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.

Instructions in long prompts: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 100.0% 100.0% to 100.0% 33 records $0.23998 run on subscription; shown at API list price for comparison 783tested in Claude Code
Claude Fable 5.1 100.0% 100.0% to 100.0% 50 records $0.16619 run on subscription; shown at API list price for comparison 362tested in Claude Code

Paired difference: +0.0 points +0.0 to +0.0, over 7 committed items both versions answered. could not be measured.

Offering the model money

On offering the model money, Claude Fable 5.1 scores 40320% unaided against 32468% for Claude Fable 5; scores higher.

Offering the model money: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 32468.0% 31347.7% to 33588.3% 50 records $0.20286 run on subscription; shown at API list price for comparison 86tested in Claude Code
Claude Fable 5.1 40320.0% 38542.7% to 42097.3% 50 records $0.23338 run on subscription; shown at API list price for comparison 3tested in Claude Code

Paired difference: +7852.0 points +6133.2 to +9570.8, over 10 committed items both versions answered. scores higher.

Questions about recent events

On questions about recent events, Claude Fable 5.1 scores 92% unaided against 88% for Claude Fable 5; no measurable change.

Questions about recent events: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 87.8% 77.5% to 98.0% 30 items $0.04206 run on subscription; shown at API list price for comparison 336tested in Claude Code
Claude Fable 5.1 92.2% 83.0% to 100.0%† 30 items $0.04170 run on subscription; shown at API list price for comparison 347tested in Claude Code

Paired difference: +3.0 points −8.1 to +14.1, over 30 committed items both versions answered. no measurable change.

Getting the length right

On getting the length right, Claude Fable 5.1 scores 50% unaided against 76% for Claude Fable 5; scores lower.

Getting the length right: Claude Fable 5 against Claude Fable 5.1, unaided.
VersionUnaided scoreCost per taskThinking per taskMeasured on
Claude Fable 5 76.0% 53.7% to 98.4% 10 items $0.16668 run on subscription; shown at API list price for comparison 437tested in Claude Code
Claude Fable 5.1 50.2% 21.9% to 78.5% 10 items $0.22184 run on subscription; shown at API list price for comparison 531tested in Claude Code

Paired difference: −26.4 points −50.2 to −2.7, over 10 committed items both versions answered. scores lower.

This is the first point in a version series, not a rate. One interval between two versions supports no curve and no extrapolation to a third. Claude Fable 5 and Claude Fable 5.1 price identically, so a cost difference between them is a token difference.

Thinking was recorded rather than refused, by prior registration

Alone among the models on this instrument, this one was registered in advance to have its thinking RECORDED rather than suppressed, and a nonzero count on it is not a gate failure. That was registered before the run precisely so that a result on this model could not later be explained away, or quietly discarded, on the grounds that it thought. The figures beside this are a measurement; the zeroes on the other models are a refusal holding, and the two are not the same kind of number.

Records reporting a nonzero count
492 of 1330
Thinking tokens recorded
143,447

How this model was called, tested in Claude Code

On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this instrument returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.

These cells are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.

Every figure on this page is computed at build time from the committed records. Nothing here averages a figure from one instrument with a figure from another, and the page carries no spread across them, because a model measured two ways has two results and not one.