Model under test
Claude Opus 5.5
Every result this site holds for Claude Opus 5.5. We tested it inside Claude Code.
What was measured, and where
We ran 37 separate tests on this model, listed by test method below and never added together. How the test methods were compared.
Prompting tips, in Claude Code, each task asked once
12 tests, tested in Claude Code, Opus 5.5 series, asked once.
- No measured effect 6 of 12
- Could not measure 5 of 12
- Holds 1 of 12
Skills, in Claude Code: Opus 5.5 series, wider run, three ways
25 tests, tested in Claude Code, Opus 5.5 series, wider run, three ways.
- No measured effect 20 of 25
- Could not measure 5 of 25
What these figures can be set beside
Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.
- Measured on
- tested in Claude Code.
- Not directly comparable
- Gemini 3.1 Flash Lite; GPT-5 mini; GPT-5.4 mini were tested through the API, which this model did not run. No figure below is theirs and none of theirs is set against these.
- Not directly comparable
- GPT-5.4 mini; GPT-5.6 Luna; GPT-5.6 Terra were tested inside Codex, which this model did not run. No figure below is theirs and none of theirs is set against these.
The model measured on both, job by job
Claude Haiku 4.5 was tested through the API and inside Claude Code. Its readings are the only thing connecting them, and they are stated job by job, never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.
Show the readings, job by job (17 jobs)
- Accessibility audit Claude Haiku 4.5 scored 40.0% (26.3% to 53.7%) tested via API, First run, with and without the skill, and 45.5% (35.0% to 56.0%) tested in Claude Code, Second run, every task twice, and 46.1% (34.1% to 58.1%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Code review Claude Haiku 4.5 scored 58.3% (47.8% to 68.7%) tested in Claude Code, Wider run, three ways. That is one test method only, so it bridges nothing on its own.
- On-page audit Claude Haiku 4.5 scored 23.9% (1.8% to 46.1%) tested via API, First run, with and without the skill, and 23.1% (6.6% to 39.6%) tested in Claude Code, Second run, every task twice, and 23.1% (3.6% to 42.6%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Skill authoring Claude Haiku 4.5 scored 61.8% (52.3% to 71.3%) tested via API, First run, with and without the skill, and 70.0% (63.5% to 76.5%) tested in Claude Code, Second run, every task twice, and 67.3% (59.7% to 74.9%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Spec writing Claude Haiku 4.5 scored 99.5% (98.7% to 100.0%†) tested via API, First run, with and without the skill, and 98.6% (97.7% to 99.6%) tested in Claude Code, Second run, every task twice, and 99.2% (98.7% to 99.7%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Getting clean JSON back Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 0.0% (0.0% to 0.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Instructions before or after Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Using tags and formatting Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Showing examples Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested in Claude Code. That is one test method only, so it bridges nothing on its own.
- Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Better step-by-step answers Claude Haiku 4.5 scored 30.0% (17.2% to 42.8%) tested via API, Main run, tip by tip, and 20.0% (8.8% to 31.2%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Stop made-up answers Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 22.0% (10.4% to 33.6%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Instructions in long prompts Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Offering the model money Claude Haiku 4.5 scored 204 words (199 words to 209 words) tested via API, Main run, tip by tip, and 241 words (236 words to 247 words) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Questions about recent events Claude Haiku 4.5 scored 85.6% (73.2% to 98.0%) tested via API, Main run, tip by tip, and 96.7% (90.1% to 100.0%†) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Getting the length right Claude Haiku 4.5 scored 24.7% (1.2% to 48.2%) tested via API, Main run, tip by tip, and 27.5% (0.0% to 54.9%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
- The grader, separately
- a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.
What this model does on its own, one job at a time
We measured this model on 17 kinds of work, and the scores for every one of them are below.
Show the figures, one job at a time
One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.
There is no cost-against-score chart for this model. This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess.
Accessibility audit · seeded accessibility defects
tested in Claude Code, Opus 5.5 series, wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 85.8% 74.4% to 97.2% 40 records over 10 items, pooled from 4 entries | 86.9% alirezarezvani-a11y-audit (of 4 entries) | not measured | not measured in, 562 out | 472 | one pass, not measured | 18.3 s 16.7 s to 19.6 s across the middle half of 120 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Code review · seeded code defects
tested in Claude Code, Opus 5.5 series, wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 85.4% 78.3% to 92.5% 80 records over 10 items, pooled from 8 entries | 88.2% tied: everyinc-ce-code-review, open-gsd-code-review (of 8 entries) | not measured | not measured in, 311 out | 232 | one pass, not measured | 15.0 s 13.4 s to 16.7 s across the middle half of 242 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 80 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 80 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
On-page audit · seeded on-page issues
tested in Claude Code, Opus 5.5 series, wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 93.3% 86.1% to 100.0%† 30 records over 10 items, pooled from 3 entries | 98.3% rampstackco-seo-onpage (of 3 entries) | not measured | not measured in, 401 out | 334 | one pass, not measured | 15.3 s 14.4 s to 16.2 s across the middle half of 90 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
A cost per task needs both sides of the call. Only 0 of the 30 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 30 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Skill authoring · specification requirements
tested in Claude Code, Opus 5.5 series, wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 99.8% 99.3% to 100.0%† 40 records over 10 items, pooled from 4 entries | 100.0% tied: anthropics-skill-creator, openclaw-skill-creator, rampstackco-skill-creation-walkthrough, yusufkaraaslan-skill-builder (of 4 entries) | not measured | not measured in, 3,531 out | 124 | one pass, not measured | 52.8 s 43.8 s to 67.1 s across the middle half of 120 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
A cost per task needs both sides of the call. Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 40 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Spec writing · required sections and acceptance criteria
tested in Claude Code, Opus 5.5 series, wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 99.5% 98.9% to 100.0%† 60 records over 10 items, pooled from 6 entries | 100.0% tied: affaan-m-product-capability, github-prd, phuryn-create-prd, rampstackco-pm-spec-writing (of 6 entries) | not measured | not measured in, 2,271 out | 56 | one pass, not measured | 36.0 s 32.9 s to 49.2 s across the middle half of 180 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
A cost per task needs both sides of the call. Only 0 of the 60 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 60 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting clean JSON back
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 0.0% 0.0% to 0.0% 10 items | 100.0% C01-json-schema | not measured | not measured in, 115 out | 0 | one pass, not measured | 12.5 s 12.1 s to 13.6 s across the middle half of 40 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions before or after
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 100.0% 100.0% to 100.0% 10 items | 100.0% C02-instruction-after | not measured | not measured in, 23 out | 16 | one pass, not measured | 11.7 s 11.6 s to 12.1 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Using tags and formatting
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 100.0% 100.0% to 100.0% 10 items | 100.0% C03-xml-delimiters | not measured | not measured in, 484 out | 149 | one pass, not measured | 16.3 s 15.5 s to 17.2 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 100.0% 100.0% to 100.0% 10 items | 100.0% C04-three-shot-format | not measured | not measured in, 29 out | 0 | one pass, not measured | 11.9 s 11.8 s to 12.1 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 82.9% 79.1% to 86.6% 10 items | 100.0% C04b-zero-vs-one | not measured | not measured in, 147 out | 108 | one pass, not measured | 12.8 s 12.4 s to 13.1 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 82.9% 79.1% to 86.6% 10 items | 100.0% C04b-zero-vs-three | not measured | not measured in, 148 out | 110 | one pass, not measured | 13.1 s 12.7 s to 13.3 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Better step-by-step answers
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 40.0% 8.0% to 72.0% 10 items | 50.0% C05-think-step-by-step | not measured | not measured in, 15 out | 3 | one pass, not measured | 12.1 s 11.9 s to 12.2 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Stop made-up answers
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 90.0% 70.4% to 100.0%† 10 items | 100.0% C06-permit-idk | not measured | not measured in, 89 out | 13 | one pass, not measured | 12.2 s 12.0 s to 13.0 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions in long prompts
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 100.0% 100.0% to 100.0% 10 items | 100.0% C07-instruction-at-end | not measured | not measured in, 414 out | 107 | one pass, not measured | 15.0 s 14.4 s to 16.2 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Offering the model money
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Answer length, on its own | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 508 words 466 to 549 10 items | 633 words C08-tip-length | not measured | not measured in, 1,176 out | 31 | one pass, not measured | 23.9 s 23.4 s to 25.7 s across the middle half of 20 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Questions about recent events
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 100.0% 100.0% to 100.0% 30 items | 93.3% C16-cutoff-disclosure | not measured | not measured in, 283 out | 161 | one pass, not measured | 14.1 s 13.7 s to 14.5 s across the middle half of 60 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 30 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 30 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting the length right
tested in Claude Code, Opus 5.5 series, each task asked once
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task | Same answer on a repeat | Time per call |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 80.1% 62.5% to 97.8% 10 items | 100.0% C17-exact-length | not measured | not measured in, 522 out | 163 | one pass, not measured | 16.3 s 15.0 s to 19.9 s across the middle half of 40 calls |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
A cost per task needs both sides of the call. Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set.
Only 0 of the 10 runs behind this figure reported an input total, so there is no way to total it up. The figure is left blank rather than guessed from the runs that did report it: those runs are not a random half of the set. Why some figures are blank.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Opus 5.5 beside Fable 5.1, and after Opus 5
Beside Fable 5.1
The maker says Opus 5.5 comes close to Fable 5.1. Both models think on every answer, so this is the fair pair. We measured one thing: whether the two likely ranges overlap on each job. Overlap means we cannot tell them apart here. It does not mean they are the same.
Prompting tips: the two likely ranges overlap on 14 of 14 jobs.
Every job, with its figures
| Job | Claude Opus 5.5 | Claude Fable 5.1 | Overlap | Cost per task |
|---|---|---|---|---|
| Getting clean JSON back | +100.0 [+100.0, +100.0] | +100.0 [+100.0, +100.0] | yes | $0.0045 / $0.0086 |
| Instructions before or after | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0035 / $0.0053 |
| Using tags and formatting | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0123 / $0.0355 |
| Showing examples | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0033 / $0.0062 |
| Better step-by-step answers | +10.0 [-9.6, +29.6] | +10.0 [-9.6, +29.6] | yes | $0.0030 / $0.0058 |
| Stop made-up answers | +10.0 [-9.6, +29.6] | +30.0 [+0.1, +59.9] | yes | $0.0030 / $0.0049 |
| Instructions in long prompts | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0115 / $0.0273 |
| Offering the model money | +125.5 words [+61.5 words, +189.5 words] | +58.8 words [+15.5 words, +102.1 words] | yes | $0.0295 / $0.0525 |
| Showing examples | +17.1 [+13.4, +20.9] | +17.1 [+13.4, +20.9] | yes | $0.0052 / $0.0086 |
| Showing examples | +17.1 [+13.4, +20.9] | +18.6 [+14.3, +22.8] | yes | $0.0049 / $0.0083 |
| Getting the length right | +19.9 [+2.2, +37.5] | +52.2 [+21.8, +82.5] | yes | $0.0215 / $0.0503 |
| Questions about recent events | -6.7 [-15.7, +2.4] | -10.0 [-20.9, +0.9] | yes | $0.0076 / $0.0165 |
| Getting clean JSON back, harder set | +100.0 [+100.0, +100.0] | +100.0 [+100.0, +100.0] | yes | $0.0071 / $0.0161 |
| Getting the length right, harder set | +99.8 [+99.3, +100.2] | +99.8 [+99.3, +100.2] | yes | $0.0157 / $0.0364 |
Cost is what each side's own answers would have cost at the maker's list price. Nobody was charged it: every answer ran on a subscription.
Skills: the two likely ranges overlap on 22 of 25 jobs.
Every job, with its figures
| Job | Claude Opus 5.5 | Claude Fable 5.1 | Overlap | Cost per task |
|---|---|---|---|---|
S07-seo-onpage | +4.2 [-4.7, +13.0] | +5.4 [-1.7, +12.4] | yes | $0.0461 / $0.1273 |
S08-ecc-seo | -3.3 [-9.9, +3.2] | +6.2 [-7.4, +19.8] | yes | $0.0188 / $0.0585 |
S20-alirezarezvani-seo-audit | +5.0 [-4.8, +14.8] | +0.0 [+0.0, +0.0] | yes | $0.0465 / $0.1269 |
S03-accessibility-audit | +0.0 [+0.0, +0.0] | +4.0 [-1.2, +9.2] | yes | $0.0881 / $0.2569 |
S04-ecc-accessibility | +2.1 [-5.1, +9.3] | -6.0 [-14.4, +2.4] | yes | $0.0266 / $0.1000 |
S14-wshobson-wcag-audit-patterns | -2.8 [-11.2, +5.7] | -2.0 [-5.9, +1.9] | yes | $0.0391 / $0.1331 |
S15-alirezarezvani-a11y-audit | -0.7 [-5.6, +4.1] | +2.0 [-7.1, +11.1] | yes | $0.1153 / $0.3322 |
S01-skill-creation-walkthrough | +0.0 [+0.0, +0.0] | -14.5 [-33.6, +4.5] | yes | $0.1230 / $0.3251 |
S02-skill-creator | +0.0 [+0.0, +0.0] | -43.6 [-66.9, -20.4] | no | $0.1287 / $0.4005 |
S12-openclaw-skill-creator | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0470 / $0.1088 |
S13-yusufkaraaslan-skill-builder | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0749 / $0.1633 |
S05-pm-spec-writing | +0.0 [+0.0, +0.0] | +3.6 [+0.7, +6.5] | no | $0.0912 / $0.2096 |
S06-product-capability | +1.8 [-0.6, +4.2] | +0.9 [-2.3, +4.1] | yes | $0.0553 / $0.1280 |
S16-phuryn-create-prd | +0.9 [-0.9, +2.7] | +3.6 [+0.7, +6.5] | yes | $0.0529 / $0.1146 |
S17-alirezarezvani-spec-driven-workflow | -0.9 [-4.1, +2.3] | -13.6 [-16.6, -10.7] | no | $0.1463 / $0.4132 |
S18-addyosmani-spec-driven-development | +0.9 [-2.3, +4.1] | +0.0 [-2.7, +2.7] | yes | $0.0649 / $0.1590 |
S19-github-prd | +1.8 [-0.6, +4.2] | -0.9 [-5.1, +3.2] | yes | $0.0584 / $0.1462 |
S11-code-review-web | +2.0 [-1.9, +5.9] | +0.0 [+0.0, +0.0] | yes | $0.0606 / $0.1734 |
S22-mattpocock-code-review | -3.7 [-8.5, +1.1] | +1.7 [-1.6, +4.9] | yes | $0.0233 / $0.0764 |
S23-alirezarezvani-pr-review-expert | +1.7 [-1.6, +4.9] | -1.7 [-4.9, +1.6] | yes | $0.0361 / $0.1123 |
S24-addyosmani-code-review-and-quality | -3.3 [-7.7, +1.0] | +0.0 [+0.0, +0.0] | yes | $0.0412 / $0.1235 |
S25-jeffallan-code-reviewer | -1.7 [-4.9, +1.6] | +0.0 [+0.0, +0.0] | yes | $0.0634 / $0.1846 |
S26-everyinc-ce-code-review | +3.3 [-3.2, +9.9] | +0.8 [-5.4, +7.0] | yes | $0.4647 / $1.1956 |
S27-wshobson-code-review-excellence | +0.0 [+0.0, +0.0] | -5.7 [-11.3, +0.0] | yes | $0.0343 / $0.1052 |
S28-open-gsd-code-review | +3.3 [-3.2, +9.9] | +0.0 [+0.0, +0.0] | yes | $0.0165 / $0.0642 |
Cost is what each side's own answers would have cost at the maker's list price. Nobody was charged it: every answer ran on a subscription.
After Opus 5
Opus 5 was tested with its thinking switched off, and the switch held on every answer. Opus 5.5 cannot switch its thinking off. So this pair sets a model that did not think beside one that always does. It is the first pair like this we have published. No difference in it can be put down to either model.
Prompting tips: the two likely ranges overlap on 12 of 14 jobs.
Every job, with its figures
| Job | Claude Opus 5.5 | Claude Opus 5 | Overlap | Cost per task |
|---|---|---|---|---|
| Getting clean JSON back | +100.0 [+100.0, +100.0] | +100.0 [+100.0, +100.0] | yes | $0.0045 / $0.0044 |
| Instructions before or after | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0035 / $0.0026 |
| Using tags and formatting | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0123 / $0.0093 |
| Showing examples | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0033 / $0.0031 |
| Better step-by-step answers | +10.0 [-9.6, +29.6] | +0.0 [+0.0, +0.0] | yes | $0.0030 / $0.0028 |
| Stop made-up answers | +10.0 [-9.6, +29.6] | +50.0 [+17.3, +82.7] | yes | $0.0030 / $0.0032 |
| Instructions in long prompts | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0115 / $0.0163 |
| Offering the model money | +125.5 words [+61.5 words, +189.5 words] | +264.8 words [+150.3 words, +379.3 words] | yes | $0.0295 / $0.0381 |
| Showing examples | +17.1 [+13.4, +20.9] | +18.6 [+14.3, +22.8] | yes | $0.0052 / $0.0038 |
| Showing examples | +17.1 [+13.4, +20.9] | +15.7 [+12.9, +18.5] | yes | $0.0049 / $0.0034 |
| Getting the length right | +19.9 [+2.2, +37.5] | +81.1 [+58.8, +103.4] | no | $0.0215 / $0.0155 |
| Questions about recent events | -6.7 [-15.7, +2.4] | +6.7 [-9.4, +22.8] | yes | $0.0076 / $0.0056 |
| Getting clean JSON back, harder set | +100.0 [+100.0, +100.0] | +100.0 [+100.0, +100.0] | yes | $0.0071 / $0.0077 |
| Getting the length right, harder set | +99.8 [+99.3, +100.2] | +93.1 [+90.2, +95.9] | no | $0.0157 / $0.0111 |
Cost is what each side's own answers would have cost at the maker's list price. Nobody was charged it: every answer ran on a subscription.
Skills: the two likely ranges overlap on 25 of 25 jobs.
Every job, with its figures
| Job | Claude Opus 5.5 | Claude Opus 5 | Overlap | Cost per task |
|---|---|---|---|---|
S07-seo-onpage | +4.2 [-4.7, +13.0] | -4.2 [-17.8, +9.4] | yes | $0.0461 / $0.0483 |
S08-ecc-seo | -3.3 [-9.9, +3.2] | -7.9 [-16.1, +0.4] | yes | $0.0188 / $0.0150 |
S20-alirezarezvani-seo-audit | +5.0 [-4.8, +14.8] | -3.3 [-14.3, +7.6] | yes | $0.0465 / $0.0481 |
S03-accessibility-audit | +0.0 [+0.0, +0.0] | -8.0 [-20.0, +4.0] | yes | $0.0881 / $0.0907 |
S04-ecc-accessibility | +2.1 [-5.1, +9.3] | +3.8 [-5.4, +13.0] | yes | $0.0266 / $0.0201 |
S14-wshobson-wcag-audit-patterns | -2.8 [-11.2, +5.7] | -6.7 [-21.4, +8.0] | yes | $0.0391 / $0.0335 |
S15-alirezarezvani-a11y-audit | -0.7 [-5.6, +4.1] | +0.8 [-7.8, +9.4] | yes | $0.1153 / $0.1275 |
S01-skill-creation-walkthrough | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.1230 / $0.1379 |
S02-skill-creator | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.1287 / $0.1509 |
S12-openclaw-skill-creator | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0470 / $0.0412 |
S13-yusufkaraaslan-skill-builder | +0.0 [+0.0, +0.0] | +0.0 [+0.0, +0.0] | yes | $0.0749 / $0.0756 |
S05-pm-spec-writing | +0.0 [+0.0, +0.0] | +2.7 [-1.9, +7.4] | yes | $0.0912 / $0.1039 |
S06-product-capability | +1.8 [-0.6, +4.2] | +1.8 [-1.7, +5.4] | yes | $0.0553 / $0.0745 |
S16-phuryn-create-prd | +0.9 [-0.9, +2.7] | +3.6 [+0.7, +6.5] | yes | $0.0529 / $0.0680 |
S17-alirezarezvani-spec-driven-workflow | -0.9 [-4.1, +2.3] | -0.9 [-4.1, +2.3] | yes | $0.1463 / $0.1704 |
S18-addyosmani-spec-driven-development | +0.9 [-2.3, +4.1] | +0.0 [+0.0, +0.0] | yes | $0.0649 / $0.0784 |
S19-github-prd | +1.8 [-0.6, +4.2] | +0.0 [+0.0, +0.0] | yes | $0.0584 / $0.0805 |
S11-code-review-web | +2.0 [-1.9, +5.9] | +0.0 [+0.0, +0.0] | yes | $0.0606 / $0.0680 |
S22-mattpocock-code-review | -3.7 [-8.5, +1.1] | +2.0 [-1.9, +5.9] | yes | $0.0233 / $0.0207 |
S23-alirezarezvani-pr-review-expert | +1.7 [-1.6, +4.9] | -1.7 [-4.9, +1.6] | yes | $0.0361 / $0.0373 |
S24-addyosmani-code-review-and-quality | -3.3 [-7.7, +1.0] | +2.0 [-1.9, +5.9] | yes | $0.0412 / $0.0442 |
S25-jeffallan-code-reviewer | -1.7 [-4.9, +1.6] | +2.0 [-1.9, +5.9] | yes | $0.0634 / $0.0701 |
S26-everyinc-ce-code-review | +3.3 [-3.2, +9.9] | +0.0 [+0.0, +0.0] | yes | $0.4647 / $0.5750 |
S27-wshobson-code-review-excellence | +0.0 [+0.0, +0.0] | +2.0 [-1.9, +5.9] | yes | $0.0343 / $0.0349 |
S28-open-gsd-code-review | +3.3 [-3.2, +9.9] | +0.0 [+0.0, +0.0] | yes | $0.0165 / $0.0147 |
Cost is what each side's own answers would have cost at the maker's list price. Nobody was charged it: every answer ran on a subscription.
This model was also run through Claude Code, on a path that bills nothing. How it was called, and what that means for every cost figure on this page, is below.
Show the full breakdown, tested in Claude Code
How this model was called, tested in Claude Code
On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this path returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.
These results are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.