Claude Opus vs Sonnet vs Haiku, Tested

Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5 did the same work on eight jobs. On 7 of them the scores could be ranked, and on 3 of those they could not be told apart.

Job by job

A solid mark is the higher score. “Too close to call” means the gap is smaller than the scores would move on another try.

In Claude Code

JobClaude Opus 5Claude Sonnet 5Claude Haiku 4.5
Code review tier 193.5% best $0.00948 a task same answer twice: one pass, not measured time: not timed74.9% too close to call $0.00406 a task same answer twice: one pass, not measured time: not timed58.3% $0.00120 a task same answer twice: one pass, not measured time: not timed
On-page audit tier 188.1% best $0.00715 a task same answer twice: one pass, not measured time: not timed53.0% $0.00362 a task same answer twice: one pass, not measured time: not timed23.1% $0.00105 a task same answer twice: one pass, not measured time: not timed
Skill authoring tier 1100.0% best $0.05962 a task same answer twice: one pass, not measured time: not timed100.0% best $0.02027 a task same answer twice: one pass, not measured time: not timed67.3% $0.00504 a task same answer twice: one pass, not measured time: not timed
Spec writing tier 192.3% $0.06098 a task same answer twice: one pass, not measured time: not timed96.2% $0.01601 a task same answer twice: one pass, not measured time: not timed99.2% best $0.00315 a task same answer twice: one pass, not measured time: not timed
Getting clean JSON back tier 20.0% no cost published same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times0.0% no cost published same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times0.0% no cost published same answer twice: 10 of 10 time: 11.4 s a call, across every tip asked five times
Getting the length right tier 20.3% too close to call no cost published same answer twice: 8 of 10 time: 12.0 s a call, across every tip asked five times1.4% too close to call no cost published same answer twice: 7 of 10 time: 11.8 s a call, across every tip asked five times4.2% best no cost published same answer twice: 7 of 10 time: 11.4 s a call, across every tip asked five times
Questions about recent events tier 180.3% too close to call $0.01566 a task same answer twice: 23 of 30 time: 12.0 s a call, across every tip asked five times93.3% too close to call $0.00757 a task same answer twice: 22 of 30 time: 11.8 s a call, across every tip asked five times96.7% best $0.00220 a task same answer twice: 28 of 30 time: 11.4 s a call, across every tip asked five times
Better step-by-step answers tier 140.0% best $0.00807 a task same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times40.0% best $0.00630 a task same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times20.0% too close to call $0.00135 a task same answer twice: 8 of 10 time: 11.4 s a call, across every tip asked five times

Where they differ

  • On code review, tested inside Claude Code, Claude Opus 5 (93.5%) and Claude Sonnet 5 (74.9%) scored higher than Claude Haiku 4.5 (58.3%).
  • On on-page audit, tested inside Claude Code, Claude Opus 5 (88.1%) scored higher than Claude Sonnet 5 (53.0%) and Claude Haiku 4.5 (23.1%).
  • On skill authoring, tested inside Claude Code, Claude Opus 5 (100.0%) and Claude Sonnet 5 (100.0%) scored higher than Claude Haiku 4.5 (67.3%).
  • On spec writing, tested inside Claude Code, Claude Haiku 4.5 (99.2%) scored higher than Claude Opus 5 (92.3%) and Claude Sonnet 5 (96.2%).

When an input was missing

One input was left out on purpose. In one run nothing said so. In the other, the model was told.

JobClaude Opus 5, not toldClaude Sonnet 5, not toldClaude Haiku 4.5, not toldClaude Opus 5, toldClaude Sonnet 5, toldClaude Haiku 4.5, told
Link graph and metadata parity auditwent ahead 0 of 5went ahead 5 of 5went ahead 5 of 5went ahead 0 of 5went ahead 0 of 5not run
Corpus integrity and correctionwent ahead 0 of 5went ahead 0 of 5went ahead 2 of 4went ahead 0 of 5went ahead 0 of 5not run
Traffic drop triagewent ahead 0 of 5went ahead 0 of 5went ahead 1 of 4went ahead 0 of 5went ahead 0 of 5not run

No version pair on this site links these models. Every figure here was tested inside Claude Code.

How this comparison was made

Two models are compared on a job only where they answered exactly the same tasks, on the same test method, at the same level of difficulty, the same number of times. A page like this one exists only where that holds on at least three jobs. Eight jobs qualified here.

  • Code review, in claude code: read from expansion-cohort.
  • On-page audit, in claude code: read from expansion-cohort.
  • Skill authoring, in claude code: read from expansion-cohort.
  • Spec writing, in claude code: read from expansion-cohort.
  • Getting clean JSON back, in claude code: read from tip-tier2-run.
  • Getting the length right, in claude code: read from tip-tier2-run.
  • Questions about recent events, in claude code: read from launch.
  • Better step-by-step answers, in claude code: read from launch.

The same figures as data: /compare/claude-opus-vs-sonnet-vs-haiku.json.