Comparing models on identical work
Models are compared only on identical work: the same task set, on the same instrument, in the same run. There is no table on this page that mixes two of any of those, and no figure anywhere that was averaged across them.
That is why there are many small tables rather than one large one. A single ranking across task sets would be a number that describes no measurement: the task sets have different answer keys, different item counts and different scorers.
Every figure here is derived from committed records. Nothing on this page was written by hand, and the methodology records how each was measured.
What changed between two versions
Higher on 4 task sets, lower on 1, no measurable change on 4, could not be measured on 6.
Whether a figure moved because the version changed is not something this design can say. Neither version was run twice and thinking is not deterministic on this transport, so a moved figure and a re-run of the same version are not distinguishable here. Every difference below is reported as measured and never as explained.
On-page audit
On on-page audit, Claude Fable 5.1 finds 98% of seeded on-page issues unaided against 92% for Claude Fable 5; scores higher.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 92.3% 85.3% to 99.4% 10 items | $0.03843 run on subscription; shown at API list price for comparison | 483 | tested in Claude Code, Skill cohort (one pass) |
| Claude Fable 5.1 | 98.3% 95.1% to 100.0%† 10 items | $0.03907 run on subscription; shown at API list price for comparison | 499 | tested in Claude Code, Fable 5.1, one pass |
Paired difference: +6.0 points +1.1 to +10.9, over 10 committed items both versions answered. scores higher.
Skill authoring
On skill authoring, Claude Fable 5.1 finds 100% of specification requirements unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 10 items | $0.10060 run on subscription; shown at API list price for comparison | 64 | tested in Claude Code, Skill cohort (one pass) |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 10 items | $0.15512 run on subscription; shown at API list price for comparison | 145 | tested in Claude Code, Fable 5.1, one pass |
Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.
Spec writing
On spec writing, Claude Fable 5.1 finds 98% of required sections and acceptance criteria unaided against 93% for Claude Fable 5; scores higher.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 93.2% 91.2% to 95.2% 10 items | $0.08369 run on subscription; shown at API list price for comparison | 20 | tested in Claude Code, Skill cohort (one pass) |
| Claude Fable 5.1 | 97.7% 96.2% to 99.2% 10 items | $0.11240 run on subscription; shown at API list price for comparison | 85 | tested in Claude Code, Fable 5.1, one pass |
Paired difference: +4.5 points +2.2 to +6.8, over 10 committed items both versions answered. scores higher.
Getting clean JSON back
On getting clean JSON back, Claude Fable 5.1 scores 0% unaided against 0% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 0.0% 0.0% to 0.0% 50 records | $0.04321 run on subscription; shown at API list price for comparison | 46 | tested in Claude Code |
| Claude Fable 5.1 | 0.0% 0.0% to 0.0% 50 records | $0.04114 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.
Instructions before or after
On instructions before or after, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | $0.02645 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.02928 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.
Using tags and formatting
On using tags and formatting, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 48 records | $0.11688 run on subscription; shown at API list price for comparison | 543 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.17308 run on subscription; shown at API list price for comparison | 1,351 | tested in Claude Code |
Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.
Showing examples
On showing examples, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | $0.03131 run on subscription; shown at API list price for comparison | 132 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.02474 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. could not be measured.
Showing examples
On showing examples, Claude Fable 5.1 scores 82% unaided against 82% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 82.3% 78.6% to 86.0% 10 items | $0.08170 run on subscription; shown at API list price for comparison | 872 | tested in Claude Code |
| Claude Fable 5.1 | 81.7% 77.7% to 85.7% 10 items | $0.05013 run on subscription; shown at API list price for comparison | 252 | tested in Claude Code |
Paired difference: −0.5 points −1.4 to +0.5, over 10 committed items both versions answered. no measurable change.
Showing examples
On showing examples, Claude Fable 5.1 scores 81% unaided against 82% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 82.0% 78.2% to 85.8% 10 items | $0.07734 run on subscription; shown at API list price for comparison | 798 | tested in Claude Code |
| Claude Fable 5.1 | 81.4% 77.2% to 85.7% 10 items | $0.05010 run on subscription; shown at API list price for comparison | 252 | tested in Claude Code |
Paired difference: −0.2 points −0.7 to +0.2, over 10 committed items both versions answered. no measurable change.
Better step-by-step answers
On better step-by-step answers, Claude Fable 5.1 scores 40% unaided against 40% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 40.0% 26.3% to 53.7% 50 records | $0.02052 run on subscription; shown at API list price for comparison | 107 | tested in Claude Code |
| Claude Fable 5.1 | 40.0% 26.3% to 53.7% 50 records | $0.02103 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Paired difference: +0.0 points +0.0 to +0.0, over 10 committed items both versions answered. no measurable change.
Stop made-up answers
On stop made-up answers, Claude Fable 5.1 scores 76% unaided against 24% for Claude Fable 5; scores higher.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 24.0% 12.0% to 36.0% 50 records | $0.03501 run on subscription; shown at API list price for comparison | 44 | tested in Claude Code |
| Claude Fable 5.1 | 76.0% 64.0% to 88.0% 50 records | $0.03175 run on subscription; shown at API list price for comparison | 0 | tested in Claude Code |
Paired difference: +52.0 points +32.4 to +71.6, over 10 committed items both versions answered. scores higher.
Instructions in long prompts
On instructions in long prompts, Claude Fable 5.1 scores 100% unaided against 100% for Claude Fable 5; could not be measured.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 100.0% 100.0% to 100.0% 33 records | $0.23998 run on subscription; shown at API list price for comparison | 783 | tested in Claude Code |
| Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | $0.16619 run on subscription; shown at API list price for comparison | 362 | tested in Claude Code |
Paired difference: +0.0 points +0.0 to +0.0, over 7 committed items both versions answered. could not be measured.
Offering the model money
On offering the model money, Claude Fable 5.1 scores 40320% unaided against 32468% for Claude Fable 5; scores higher.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 32468.0% 31347.7% to 33588.3% 50 records | $0.20286 run on subscription; shown at API list price for comparison | 86 | tested in Claude Code |
| Claude Fable 5.1 | 40320.0% 38542.7% to 42097.3% 50 records | $0.23338 run on subscription; shown at API list price for comparison | 3 | tested in Claude Code |
Paired difference: +7852.0 points +6133.2 to +9570.8, over 10 committed items both versions answered. scores higher.
Questions about recent events
On questions about recent events, Claude Fable 5.1 scores 92% unaided against 88% for Claude Fable 5; no measurable change.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 87.8% 77.5% to 98.0% 30 items | $0.04206 run on subscription; shown at API list price for comparison | 336 | tested in Claude Code |
| Claude Fable 5.1 | 92.2% 83.0% to 100.0%† 30 items | $0.04170 run on subscription; shown at API list price for comparison | 347 | tested in Claude Code |
Paired difference: +3.0 points −8.1 to +14.1, over 30 committed items both versions answered. no measurable change.
Getting the length right
On getting the length right, Claude Fable 5.1 scores 50% unaided against 76% for Claude Fable 5; scores lower.
| Version | Unaided score | Cost per task | Thinking per task | Measured on |
|---|---|---|---|---|
| Claude Fable 5 | 76.0% 53.7% to 98.4% 10 items | $0.16668 run on subscription; shown at API list price for comparison | 437 | tested in Claude Code |
| Claude Fable 5.1 | 50.2% 21.9% to 78.5% 10 items | $0.22184 run on subscription; shown at API list price for comparison | 531 | tested in Claude Code |
Paired difference: −26.4 points −50.2 to −2.7, over 10 committed items both versions answered. scores lower.
This is the first point in a version series, not a rate. One interval between two versions supports no curve and no extrapolation to a third. Claude Fable 5 and Claude Fable 5.1 price identically, so a cost difference between them is a token difference.
Skill classes
Accessibility audit · seeded accessibility defects
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5 mini | 69.9% 54.2% to 85.5% 20 records over 10 items, pooled from 2 entries | 68.9% rampstackco-accessibility-audit (of 2 entries) | $0.00035 billed through the API | 515 in, 111 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 63.3% 46.0% to 80.7% 20 records over 10 items, pooled from 2 entries | 72.8% rampstackco-accessibility-audit (of 2 entries) | $0.00022 billed through the API | 528 in, 59 out | not measured |
| 3 | Claude Haiku 4.5 | 40.0% 26.3% to 53.7% 20 records over 10 items, pooled from 2 entries | 69.9% rampstackco-accessibility-audit (of 2 entries) | $0.00088 billed through the API | 600 in, 55 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 20 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Skill cohort (one pass)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 89.0% 79.7% to 98.3% 20 records over 10 items, pooled from 2 entries | 87.0% rampstackco-accessibility-audit (of 2 entries) | $0.00696 run on subscription; shown at API list price for comparison | 946 in, 89 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Panel v2 (repeat sampling)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 87.3% 78.3% to 96.4% 40 records over 10 items, pooled from 2 entries | 88.0% affaan-m-accessibility (of 2 entries) | $0.01384 run on subscription; shown at API list price for comparison | 1,892 in, 175 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Sonnet 5 | 82.3% 73.9% to 90.7% 40 records over 10 items, pooled from 2 entries | 82.3% affaan-m-accessibility (of 2 entries) | $0.01118 run on subscription; shown at API list price for comparison | 2,544 in, 237 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Haiku 4.5 | 45.5% 35.0% to 56.0% 40 records over 10 items, pooled from 2 entries | 69.0% rampstackco-accessibility-audit (of 2 entries) | $0.00206 run on subscription; shown at API list price for comparison | 1,530 in, 106 out | 0 a gate holding at zero, not a measured zero |
- Claude Fable 5: withheld. 71 of 76 unaided-arm calls on this task set were answered by a different model (Claude Opus 5). metrics_v1 section 4 withholds a cell whose records contain a substitution until the affected units are re-bought and served correctly. That rule matters most here: a substitution concentrated on the control arm inflates exactly the mean this column publishes, so a score built from the calls that happened to be served correctly would be a measurement of the platform's selection rather than of the model.
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Fable 5.1, one pass
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 85.8% 72.6% to 99.1% 20 records over 10 items, pooled from 2 entries | 93.5% rampstackco-accessibility-audit (of 2 entries) | $0.05119 run on subscription; shown at API list price for comparison | 948 in, 834 out | 755 |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
On-page audit · seeded on-page issues
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 48.8% 27.1% to 70.5% 20 records over 10 items, pooled from 2 entries | 70.7% rampstackco-seo-onpage (of 2 entries) | $0.00023 billed through the API | 599 in, 53 out | not measured |
| 2 | GPT-5 mini | 44.9% 26.7% to 63.1% 20 records over 10 items, pooled from 2 entries | 49.2% rampstackco-seo-onpage (of 2 entries) | $0.00029 billed through the API | 566 in, 75 out | not measured |
| 3 | Claude Haiku 4.5 | 23.9% 1.8% to 46.1% 20 records over 10 items, pooled from 2 entries | 47.4% rampstackco-seo-onpage (of 2 entries) | $0.00085 billed through the API | 667 in, 37 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 20 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Skill cohort (one pass)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 92.3% 85.3% to 99.4% 20 records over 10 items, pooled from 2 entries | 93.3% affaan-m-seo (of 2 entries) | $0.03843 run on subscription; shown at API list price for comparison | 1,068 in, 555 out | 483 |
| 2 | Claude Opus 5 | 87.9% 76.7% to 99.1% 20 records over 10 items, pooled from 2 entries | 83.3% rampstackco-seo-onpage (of 2 entries) | $0.00710 run on subscription; shown at API list price for comparison | 1,068 in, 71 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Panel v2 (repeat sampling)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 88.5% 77.6% to 99.5% 40 records over 10 items, pooled from 2 entries | 83.3% rampstackco-seo-onpage (of 2 entries) | not measured | not measured in, 141 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Fable 5 | 88.0% 79.2% to 96.9% 40 records over 10 items, pooled from 2 entries | 94.0% affaan-m-seo (of 2 entries) | not measured | not measured in, 1,203 out | 1,059 |
| 3 | Claude Sonnet 5 | 59.0% 46.0% to 72.0% 40 records over 10 items, pooled from 2 entries | 64.7% affaan-m-seo (of 2 entries) | not measured | not measured in, 171 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Haiku 4.5 | 23.1% 6.6% to 39.6% 40 records over 10 items, pooled from 2 entries | 40.8% rampstackco-seo-onpage (of 2 entries) | not measured | not measured in, 83 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
A cost per task needs both sides of the call. 0 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Fable 5.1, one pass
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 98.3% 95.1% to 100.0%† 20 records over 10 items, pooled from 2 entries | 95.0% tied: affaan-m-seo, rampstackco-seo-onpage (of 2 entries) | $0.03907 run on subscription; shown at API list price for comparison | 1,070 in, 568 out | 499 |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Skill authoring · specification requirements
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 87.3% 80.1% to 94.4% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00088 billed through the API | 59 in, 580 out | not measured |
| 2 | GPT-5 mini | 73.2% 60.9% to 85.5% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00297 billed through the API | 61 in, 1,478 out | not measured |
| 3 | Claude Haiku 4.5 | 61.8% 52.3% to 71.3% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00618 billed through the API | 68 in, 1,222 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 20 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Skill cohort (one pass)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 20 records over 10 items, pooled from 2 entries | 100.0% anthropics-skill-creator (of 2 entries) | $0.10060 run on subscription; shown at API list price for comparison | 313 in, 1,949 out | 64 |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.06061 run on subscription; shown at API list price for comparison | 313 in, 2,362 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Panel v2 (repeat sampling)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 22 records over 10 items, pooled from 2 entries | 100.0% anthropics-skill-creator (of 2 entries) | $0.19551 run on subscription; shown at API list price for comparison | 628 in, 3,785 out | 131 |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 30 records over 10 items, pooled from 2 entries | 100.0% rampstackco-skill-creation-walkthrough (of 2 entries) | $0.12485 run on subscription; shown at API list price for comparison | 627 in, 4,869 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 40 records over 10 items, pooled from 2 entries | 92.7% anthropics-skill-creator (of 2 entries) | $0.05962 run on subscription; shown at API list price for comparison | 1,278 in, 3,719 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Haiku 4.5 | 70.0% 63.5% to 76.5% 40 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00996 run on subscription; shown at API list price for comparison | 467 in, 1,898 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Fable 5.1, one pass
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 20 records over 10 items, pooled from 2 entries | 100.0% rampstackco-skill-creation-walkthrough (of 2 entries) | $0.15512 run on subscription; shown at API list price for comparison | 315 in, 3,039 out | 145 |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Spec writing · required sections and acceptance criteria
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 99.5% 98.7% to 100.0%† 20 records over 10 items, pooled from 2 entries | 99.1% rampstackco-pm-spec-writing (of 2 entries) | $0.00332 billed through the API | 118 in, 641 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 99.1% 97.3% to 100.0%† 20 records over 10 items, pooled from 2 entries | 99.1% tied: affaan-m-product-capability, rampstackco-pm-spec-writing (of 2 entries) | $0.00080 billed through the API | 106 in, 513 out | not measured |
| 3 | GPT-5 mini | 94.5% 92.0% to 97.1% 20 records over 10 items, pooled from 2 entries | 90.0% affaan-m-product-capability (of 2 entries) | $0.00248 billed through the API | 107 in, 1,227 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
0 of 40 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Skill cohort (one pass)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 93.2% 91.2% to 95.2% 20 records over 10 items, pooled from 2 entries | 93.6% affaan-m-product-capability (of 2 entries) | $0.08369 run on subscription; shown at API list price for comparison | 386 in, 1,597 out | 20 |
| 2 | Claude Opus 5 | 92.3% 90.9% to 93.6% 20 records over 10 items, pooled from 2 entries | 93.6% rampstackco-pm-spec-writing (of 2 entries) | $0.05944 run on subscription; shown at API list price for comparison | 386 in, 2,300 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Panel v2 (repeat sampling)
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 98.6% 97.7% to 99.6% 40 records over 10 items, pooled from 2 entries | 97.7% rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 1,145 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Fable 5 | 95.7% 94.6% to 96.7% 40 records over 10 items, pooled from 2 entries | 92.7% tied: affaan-m-product-capability, rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 3,234 out | 38 |
| 3 | Claude Sonnet 5 | 94.8% 93.0% to 96.5% 40 records over 10 items, pooled from 2 entries | 95.0% rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 2,948 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Opus 5 | 92.5% 91.3% to 93.7% 40 records over 10 items, pooled from 2 entries | 94.5% rampstackco-pm-spec-writing (of 2 entries) | not measured | not measured in, 5,009 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
26 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
30 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
32 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
34 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
A cost per task needs both sides of the call. 26 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
A cost per task needs both sides of the call. 30 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
A cost per task needs both sides of the call. 32 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
A cost per task needs both sides of the call. 34 of 40 records behind this figure could state an input total; the rest carry an uncached remainder with no counters to complete it. metrics_v1 section 11 refuses a partial sum presented as a total, so the figure is absent rather than computed over the records that happen to have it: those records are not a random half of the arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code, Fable 5.1, one pass
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 97.7% 96.2% to 99.2% 20 records over 10 items, pooled from 2 entries | 99.1% rampstackco-pm-spec-writing (of 2 entries) | $0.11240 run on subscription; shown at API list price for comparison | 388 in, 2,170 out | 85 |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Prompting tips, single-turn
Better step-by-step answers
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | not measured | 220 in, 11 out | not measured |
| 2 | Claude Haiku 4.5 | 30.0% 17.2% to 42.8% 50 records | 50.0% C05-think-step-by-step | not measured | 273 in, 25 out | not measured |
| 3 | GPT-5 mini | 16.0% 5.7% to 26.3% 50 records | 50.0% C05-think-step-by-step | not measured | 237 in, 53 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 50 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | $0.02052 run on subscription; shown at API list price for comparison | 1,444 in, 122 out | 107 |
| 1 | Claude Fable 5.1 | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | $0.02103 run on subscription; shown at API list price for comparison | 1,454 in, 130 out | 0 |
| 1 | Claude Opus 5 | 40.0% 26.3% to 53.7% 50 records | 46.0% C05-think-step-by-step | $0.00807 run on subscription; shown at API list price for comparison | 1,444 in, 34 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | $0.00945 run on subscription; shown at API list price for comparison | 3,074 in, 15 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Haiku 4.5 | 20.0% 8.8% to 31.2% 50 records | 50.0% C05-think-step-by-step | $0.00135 run on subscription; shown at API list price for comparison | 1,098 in, 50 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting clean JSON back
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | not measured | 277 in, 415 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | not measured | 219 in, 386 out | not measured |
| 1 | GPT-5 mini | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | not measured | 235 in, 393 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 50 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.04321 run on subscription; shown at API list price for comparison | 1,496 in, 565 out | 46 |
| 1 | Claude Fable 5.1 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.04114 run on subscription; shown at API list price for comparison | 1,506 in, 522 out | 0 |
| 1 | Claude Haiku 4.5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.00316 run on subscription; shown at API list price for comparison | 1,102 in, 412 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.02287 run on subscription; shown at API list price for comparison | 1,496 in, 616 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | $0.01740 run on subscription; shown at API list price for comparison | 3,126 in, 535 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting the length right
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 31.0% 5.8% to 56.1% 10 items | 99.4% C17-exact-length | not measured | 120 in, 3,398 out | not measured |
| 2 | Claude Haiku 4.5 | 24.7% 1.2% to 48.2% 10 items | 94.8% C17-exact-length | not measured | 164 in, 2,238 out | not measured |
| 3 | GPT-5 mini | 1.6% 0.0% to 4.6%† 10 items | 96.8% C17-exact-length | not measured | 139 in, 5,014 out | not measured |
† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
0 of 70 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 76.0% 53.7% to 98.4% 10 items | 99.5% C17-exact-length | $0.16668 run on subscription; shown at API list price for comparison | 1,768 in, 2,980 out | 437 |
| 2 | Claude Fable 5.1 | 50.2% 21.9% to 78.5% 10 items | 99.5% C17-exact-length | $0.22184 run on subscription; shown at API list price for comparison | 1,786 in, 4,080 out | 531 |
| 3 | Claude Haiku 4.5 | 27.5% 0.0% to 54.9% 10 items | 92.5% C17-exact-length | $0.01432 run on subscription; shown at API list price for comparison | 1,319 in, 2,599 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Sonnet 5 | 17.8% 0.0% to 41.0%† 10 items | 95.2% C17-exact-length | $0.09689 run on subscription; shown at API list price for comparison | 4,050 in, 5,649 out | 0 a gate holding at zero, not a measured zero |
| 5 | Claude Opus 5 | 15.0% 0.0% to 35.1%† 10 items | 97.3% C17-exact-length | $0.18907 run on subscription; shown at API list price for comparison | 1,768 in, 7,209 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions before or after
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | not measured | 981 in, 41 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | not measured | 986 in, 41 out | not measured |
| 1 | GPT-5 mini | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | not measured | 900 in, 71 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 50 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.02645 run on subscription; shown at API list price for comparison | 2,465 in, 36 out | 0 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.02928 run on subscription; shown at API list price for comparison | 2,748 in, 36 out | 0 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.00201 run on subscription; shown at API list price for comparison | 1,806 in, 41 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.01323 run on subscription; shown at API list price for comparison | 2,465 in, 36 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | $0.01282 run on subscription; shown at API list price for comparison | 4,095 in, 36 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions in long prompts
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | not measured | 1,386 in, 878 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | not measured | 1,249 in, 398 out | not measured |
| 1 | GPT-5 mini | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | not measured | 1,274 in, 845 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 50 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 33 records | 100.0% C07-instruction-at-end | $0.23998 run on subscription; shown at API list price for comparison | 4,267 in, 3,946 out | 783 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.16619 run on subscription; shown at API list price for comparison | 6,464 in, 2,031 out | 362 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.00664 run on subscription; shown at API list price for comparison | 2,211 in, 885 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.08287 run on subscription; shown at API list price for comparison | 3,170 in, 2,681 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | $0.03215 run on subscription; shown at API list price for comparison | 4,800 in, 1,183 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Offering the model money
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 51850.0% 50569.6% to 53130.4% 50 records | 57500.0% C08-tip-length | not measured | 59 in, 3,513 out | not measured |
| 2 | GPT-5 mini | 35474.0% 33413.0% to 37535.0% 50 records | 53434.0% C08-tip-length | not measured | 90 in, 2,337 out | not measured |
| 3 | Claude Haiku 4.5 | 20408.0% 19916.2% to 20899.8% 50 records | 24824.0% C08-tip-length | not measured | 103 in, 1,619 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 50 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 54422.0% 50905.4% to 57938.6% 50 records | 82768.0% C08-tip-length | $0.15452 run on subscription; shown at API list price for comparison | 1,247 in, 5,931 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Fable 5.1 | 40320.0% 38542.7% to 42097.3% 50 records | 50828.0% C08-tip-length | $0.23338 run on subscription; shown at API list price for comparison | 1,257 in, 4,416 out | 3 |
| 3 | Claude Sonnet 5 | 37712.0% 36224.5% to 39199.5% 50 records | 45708.0% C08-tip-length | $0.07162 run on subscription; shown at API list price for comparison | 2,877 in, 4,199 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Fable 5 | 32468.0% 31347.7% to 33588.3% 50 records | 38670.0% C08-tip-length | $0.20286 run on subscription; shown at API list price for comparison | 1,247 in, 3,808 out | 86 |
| 5 | Claude Haiku 4.5 | 24138.0% 23624.8% to 24651.2% 50 records | 25800.0% C08-tip-length | $0.01031 run on subscription; shown at API list price for comparison | 928 in, 1,877 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Questions about recent events
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 85.6% 73.2% to 98.0% 30 items | 62.2% C16-cutoff-disclosure | not measured | 58 in, 162 out | not measured |
| 2 | GPT-5 mini | 48.6% 31.8% to 65.4% 30 items | 89.7% C16-cutoff-disclosure | not measured | 52 in, 131 out | not measured |
| 3 | Gemini 3.1 Flash Lite | 3.3% 0.0% to 9.9%† 30 items | 33.3% C16-cutoff-disclosure | not measured | 44 in, 90 out | not measured |
† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
0 of 108 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 96.7% 90.1% to 100.0%† 30 items | 99.2% C16-cutoff-disclosure | $0.00220 run on subscription; shown at API list price for comparison | 474 in, 345 out | 0 a gate holding at zero, not a measured zero |
| 2 | Claude Sonnet 5 | 93.3% 84.3% to 100.0%† 30 items | 92.5% C16-cutoff-disclosure | $0.01136 run on subscription; shown at API list price for comparison | 1,448 in, 467 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Fable 5.1 | 92.2% 83.0% to 100.0%† 30 items | 89.2% C16-cutoff-disclosure | $0.04170 run on subscription; shown at API list price for comparison | 638 in, 706 out | 347 |
| 4 | Claude Fable 5 | 87.8% 77.5% to 98.0% 30 items | 95.6% C16-cutoff-disclosure | $0.04206 run on subscription; shown at API list price for comparison | 633 in, 715 out | 336 |
| 5 | Claude Opus 5 | 80.3% 67.4% to 93.1% 30 items | 86.7% C16-cutoff-disclosure | $0.01566 run on subscription; shown at API list price for comparison | 633 in, 500 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
† The interval runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
Showing examples
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.03131 run on subscription; shown at API list price for comparison | 1,719 in, 283 out | 132 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.02474 run on subscription; shown at API list price for comparison | 1,729 in, 149 out | 0 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.00187 run on subscription; shown at API list price for comparison | 1,281 in, 118 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.01238 run on subscription; shown at API list price for comparison | 1,719 in, 151 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C04-three-shot-format | $0.01233 run on subscription; shown at API list price for comparison | 3,349 in, 152 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 97.1% C04b-zero-vs-one | not measured | 971 in, 221 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 80.0% 75.4% to 84.6% 10 items | 95.7% C04b-zero-vs-one | not measured | 871 in, 199 out | not measured |
| 3 | GPT-5 mini | 78.3% 73.0% to 83.6% 10 items | 96.9% C04b-zero-vs-one | not measured | 875 in, 218 out | not measured |
0 of 60 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 82.3% 78.6% to 86.0% 10 items | 99.7% C04b-zero-vs-one | $0.08170 run on subscription; shown at API list price for comparison | 2,621 in, 1,110 out | 872 |
| 2 | Claude Fable 5.1 | 81.7% 77.7% to 85.7% 10 items | 100.0% C04b-zero-vs-one | $0.05013 run on subscription; shown at API list price for comparison | 2,590 in, 485 out | 252 |
| 3 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 94.3% C04b-zero-vs-one | $0.00311 run on subscription; shown at API list price for comparison | 1,961 in, 230 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Opus 5 | 81.4% 77.2% to 85.7% 10 items | 98.0% C04b-zero-vs-one | $0.01882 run on subscription; shown at API list price for comparison | 2,578 in, 237 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Sonnet 5 | 81.4% 77.2% to 85.7% 10 items | 97.7% C04b-zero-vs-one | $0.01713 run on subscription; shown at API list price for comparison | 4,534 in, 235 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 100.0% C04b-zero-vs-three | not measured | 971 in, 221 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 80.0% 75.4% to 84.6% 10 items | 98.6% C04b-zero-vs-three | not measured | 871 in, 199 out | not measured |
| 3 | GPT-5 mini | 79.4% 74.3% to 84.6% 10 items | 97.7% C04b-zero-vs-three | not measured | 875 in, 219 out | not measured |
0 of 60 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 82.0% 78.2% to 85.8% 10 items | 100.0% C04b-zero-vs-three | $0.07734 run on subscription; shown at API list price for comparison | 2,578 in, 1,031 out | 798 |
| 2 | Claude Sonnet 5 | 81.7% 77.7% to 85.7% 10 items | 100.0% C04b-zero-vs-three | $0.01713 run on subscription; shown at API list price for comparison | 4,534 in, 236 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Fable 5.1 | 81.4% 77.2% to 85.7% 10 items | 100.0% C04b-zero-vs-three | $0.05010 run on subscription; shown at API list price for comparison | 2,590 in, 484 out | 252 |
| 3 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 98.9% C04b-zero-vs-three | $0.00310 run on subscription; shown at API list price for comparison | 1,961 in, 228 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Opus 5 | 81.4% 77.2% to 85.7% 10 items | 99.7% C04b-zero-vs-three | $0.01882 run on subscription; shown at API list price for comparison | 2,578 in, 237 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Stop made-up answers
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5 mini | 20.0% 8.8% to 31.2% 50 records | 100.0% C06-permit-idk | not measured | 231 in, 165 out | not measured |
| 2 | Claude Haiku 4.5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C06-permit-idk | not measured | 256 in, 271 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 0.0% 0.0% to 0.0% 50 records | 100.0% C06-permit-idk | not measured | 222 in, 87 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 50 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 76.0% 64.0% to 88.0% 50 records | 100.0% C06-permit-idk | $0.03175 run on subscription; shown at API list price for comparison | 1,450 in, 345 out | 0 |
| 2 | Claude Sonnet 5 | 38.0% 24.4% to 51.6% 50 records | 100.0% C06-permit-idk | $0.01625 run on subscription; shown at API list price for comparison | 3,070 in, 469 out | 0 a gate holding at zero, not a measured zero |
| 3 | Claude Opus 5 | 30.0% 17.2% to 42.8% 50 records | 100.0% C06-permit-idk | $0.02363 run on subscription; shown at API list price for comparison | 1,440 in, 657 out | 0 a gate holding at zero, not a measured zero |
| 4 | Claude Fable 5 | 24.0% 12.0% to 36.0% 50 records | 100.0% C06-permit-idk | $0.03501 run on subscription; shown at API list price for comparison | 1,440 in, 412 out | 44 |
| 5 | Claude Haiku 4.5 | 22.0% 10.4% to 33.6% 50 records | 100.0% C06-permit-idk | $0.00340 run on subscription; shown at API list price for comparison | 1,081 in, 463 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Using tags and formatting
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | not measured | 749 in, 759 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | not measured | 686 in, 368 out | not measured |
| 1 | GPT-5 mini | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | not measured | 708 in, 771 out | not measured |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
0 of 50 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
This run wrote no ledger line carrying a billed cost for this arm.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 100.0% 100.0% to 100.0% 48 records | 100.0% C03-xml-delimiters | $0.11688 run on subscription; shown at API list price for comparison | 2,370 in, 1,864 out | 543 |
| 1 | Claude Fable 5.1 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.17308 run on subscription; shown at API list price for comparison | 2,193 in, 3,023 out | 1,351 |
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.00533 run on subscription; shown at API list price for comparison | 1,574 in, 752 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Opus 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.04752 run on subscription; shown at API list price for comparison | 2,183 in, 1,464 out | 0 a gate holding at zero, not a measured zero |
| 1 | Claude Sonnet 5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | $0.02477 run on subscription; shown at API list price for comparison | 3,813 in, 889 out | 0 a gate holding at zero, not a measured zero |
A model on this table ran at its own sampling default rather than at the fixed setting: one vendor rejects a temperature of 0, and the Claude Code CLI exposes no sampling flag. Recorded as a divergence, not a failure.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Prompting tips, multi-turn
Prompt, then revise
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 95.7% interval not measured 10 items | 98.6% C15-iterative-refinement | not measured | not measured in, not measured out | not measured |
not measured
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 95.7% interval not measured 10 items | 98.6% C15-iterative-refinement | $0.00083 run on subscription; shown at API list price for comparison | 557 in, 55 out | not measured |
0 of 10 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Rules for the whole chat
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% interval not measured 10 items | 100.0% C13-system-prompt-placement | not measured | not measured in, not measured out | not measured |
not measured
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% interval not measured 10 items | 100.0% C13-system-prompt-placement | $0.01660 run on subscription; shown at API list price for comparison | 13,384 in, 643 out | not measured |
0 of 10 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Splitting big tasks
tested via API
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 68.7% interval not measured 10 items | 100.0% C14-prompt-chaining | not measured | not measured in, not measured out | not measured |
not measured
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
tested in Claude Code
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 68.7% interval not measured 10 items | 100.0% C14-prompt-chaining | $0.00517 run on subscription; shown at API list price for comparison | 3,902 in, 255 out | not measured |
0 of 60 records behind this figure recorded a thinking-token count. A run that recorded no count did not record zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.