Claude Opus vs Sonnet vs Haiku, Tested
Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5 did the same work on eight jobs. On 7 of them the scores could be ranked, and on 3 of those they could not be told apart.
Job by job
A solid mark is the higher score. “Too close to call” means the gap is smaller than the scores would move on another try.
In Claude Code
| Job | Claude Opus 5 | Claude Sonnet 5 | Claude Haiku 4.5 |
|---|---|---|---|
| Code review tier 1 | 93.5% best $0.00948 a task same answer twice: one pass, not measured time: not timed | 74.9% too close to call $0.00406 a task same answer twice: one pass, not measured time: not timed | 58.3% $0.00120 a task same answer twice: one pass, not measured time: not timed |
| On-page audit tier 1 | 88.1% best $0.00715 a task same answer twice: one pass, not measured time: not timed | 53.0% $0.00362 a task same answer twice: one pass, not measured time: not timed | 23.1% $0.00105 a task same answer twice: one pass, not measured time: not timed |
| Skill authoring tier 1 | 100.0% best $0.05962 a task same answer twice: one pass, not measured time: not timed | 100.0% best $0.02027 a task same answer twice: one pass, not measured time: not timed | 67.3% $0.00504 a task same answer twice: one pass, not measured time: not timed |
| Spec writing tier 1 | 92.3% $0.06098 a task same answer twice: one pass, not measured time: not timed | 96.2% $0.01601 a task same answer twice: one pass, not measured time: not timed | 99.2% best $0.00315 a task same answer twice: one pass, not measured time: not timed |
| Getting clean JSON back tier 2 | 0.0% no cost published same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times | 0.0% no cost published same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times | 0.0% no cost published same answer twice: 10 of 10 time: 11.4 s a call, across every tip asked five times |
| Getting the length right tier 2 | 0.3% too close to call no cost published same answer twice: 8 of 10 time: 12.0 s a call, across every tip asked five times | 1.4% too close to call no cost published same answer twice: 7 of 10 time: 11.8 s a call, across every tip asked five times | 4.2% best no cost published same answer twice: 7 of 10 time: 11.4 s a call, across every tip asked five times |
| Questions about recent events tier 1 | 80.3% too close to call $0.01566 a task same answer twice: 23 of 30 time: 12.0 s a call, across every tip asked five times | 93.3% too close to call $0.00757 a task same answer twice: 22 of 30 time: 11.8 s a call, across every tip asked five times | 96.7% best $0.00220 a task same answer twice: 28 of 30 time: 11.4 s a call, across every tip asked five times |
| Better step-by-step answers tier 1 | 40.0% best $0.00807 a task same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times | 40.0% best $0.00630 a task same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times | 20.0% too close to call $0.00135 a task same answer twice: 8 of 10 time: 11.4 s a call, across every tip asked five times |
Where they differ
- On code review, tested inside Claude Code, Claude Opus 5 (93.5%) and Claude Sonnet 5 (74.9%) scored higher than Claude Haiku 4.5 (58.3%).
- On on-page audit, tested inside Claude Code, Claude Opus 5 (88.1%) scored higher than Claude Sonnet 5 (53.0%) and Claude Haiku 4.5 (23.1%).
- On skill authoring, tested inside Claude Code, Claude Opus 5 (100.0%) and Claude Sonnet 5 (100.0%) scored higher than Claude Haiku 4.5 (67.3%).
- On spec writing, tested inside Claude Code, Claude Haiku 4.5 (99.2%) scored higher than Claude Opus 5 (92.3%) and Claude Sonnet 5 (96.2%).
When an input was missing
One input was left out on purpose. In one run nothing said so. In the other, the model was told.
| Job | Claude Opus 5, not told | Claude Sonnet 5, not told | Claude Haiku 4.5, not told | Claude Opus 5, told | Claude Sonnet 5, told | Claude Haiku 4.5, told |
|---|---|---|---|---|---|---|
| Link graph and metadata parity audit | went ahead 0 of 5 | went ahead 5 of 5 | went ahead 5 of 5 | went ahead 0 of 5 | went ahead 0 of 5 | not run |
| Corpus integrity and correction | went ahead 0 of 5 | went ahead 0 of 5 | went ahead 2 of 4 | went ahead 0 of 5 | went ahead 0 of 5 | not run |
| Traffic drop triage | went ahead 0 of 5 | went ahead 0 of 5 | went ahead 1 of 4 | went ahead 0 of 5 | went ahead 0 of 5 | not run |
No version pair on this site links these models. Every figure here was tested inside Claude Code.
How this comparison was made
Two models are compared on a job only where they answered exactly the same tasks, on the same test method, at the same level of difficulty, the same number of times. A page like this one exists only where that holds on at least three jobs. Eight jobs qualified here.
- Code review, in claude code: read from expansion-cohort.
- On-page audit, in claude code: read from expansion-cohort.
- Skill authoring, in claude code: read from expansion-cohort.
- Spec writing, in claude code: read from expansion-cohort.
- Getting clean JSON back, in claude code: read from tip-tier2-run.
- Getting the length right, in claude code: read from tip-tier2-run.
- Questions about recent events, in claude code: read from launch.
- Better step-by-step answers, in claude code: read from launch.
The same figures as data: /compare/claude-opus-vs-sonnet-vs-haiku.json.