Claude Sonnet vs Opus, Tested on the Same Tasks
Claude Sonnet 5 and Claude Opus 5 did the same work on 11 jobs. On 7 of them the scores could be ranked, and on 5 of those they could not be told apart.
Job by job
A solid mark is the higher score. “Too close to call” means the gap is smaller than the scores would move on another try.
In Claude Code
| Job | Claude Sonnet 5 | Claude Opus 5 |
|---|---|---|
| Code review tier 1 | 74.9% too close to call $0.00406 a task same answer twice: one pass, not measured time: not timed | 93.5% best $0.00948 a task same answer twice: one pass, not measured time: not timed |
| On-page audit tier 1 | 53.0% $0.00362 a task same answer twice: one pass, not measured time: not timed | 88.1% best $0.00715 a task same answer twice: one pass, not measured time: not timed |
| Skill authoring tier 1 | 100.0% best $0.02027 a task same answer twice: one pass, not measured time: not timed | 100.0% best $0.05962 a task same answer twice: one pass, not measured time: not timed |
| Spec writing tier 1 | 96.2% best $0.01601 a task same answer twice: one pass, not measured time: not timed | 92.3% $0.06098 a task same answer twice: one pass, not measured time: not timed |
| Link graph and metadata parity audit tier 1 | 73.7% no cost published same answer twice: not asked twice time: not timed | 92.9% best no cost published same answer twice: not asked twice time: not timed |
| Corpus integrity and correction tier 1 | 64.2% no cost published same answer twice: not asked twice time: not timed | 75.8% best no cost published same answer twice: not asked twice time: not timed |
| Traffic drop triage tier 1 | 80.0% no cost published same answer twice: not asked twice time: not timed | 100.0% best no cost published same answer twice: not asked twice time: not timed |
| Getting clean JSON back tier 2 | 0.0% no cost published same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times | 0.0% no cost published same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times |
| Getting the length right tier 2 | 1.4% best no cost published same answer twice: 7 of 10 time: 11.8 s a call, across every tip asked five times | 0.3% too close to call no cost published same answer twice: 8 of 10 time: 12.0 s a call, across every tip asked five times |
| Questions about recent events tier 1 | 93.3% best $0.00757 a task same answer twice: 22 of 30 time: 11.8 s a call, across every tip asked five times | 80.3% too close to call $0.01566 a task same answer twice: 23 of 30 time: 12.0 s a call, across every tip asked five times |
| Better step-by-step answers tier 1 | 40.0% best $0.00630 a task same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times | 40.0% best $0.00807 a task same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times |
Where they differ
- On on-page audit, tested inside Claude Code, Claude Opus 5 (88.1%) scored higher than Claude Sonnet 5 (53.0%).
- On spec writing, tested inside Claude Code, Claude Sonnet 5 (96.2%) scored higher than Claude Opus 5 (92.3%).
Gaps with no range
- On link graph and metadata parity audit, tested inside Claude Code, the scores were Claude Sonnet 5 73.7% and Claude Opus 5 92.9%. No range was published, so whether that gap is real is not tested.
- On corpus integrity and correction, tested inside Claude Code, the scores were Claude Sonnet 5 64.2% and Claude Opus 5 75.8%. No range was published, so whether that gap is real is not tested.
- On traffic drop triage, tested inside Claude Code, the scores were Claude Sonnet 5 80.0% and Claude Opus 5 100.0%. No range was published, so whether that gap is real is not tested.
When an input was missing
One input was left out on purpose. In one run nothing said so. In the other, the model was told.
| Job | Claude Sonnet 5, not told | Claude Opus 5, not told | Claude Sonnet 5, told | Claude Opus 5, told |
|---|---|---|---|---|
| Link graph and metadata parity audit | went ahead 5 of 5 | went ahead 0 of 5 | went ahead 0 of 5 | went ahead 0 of 5 |
| Corpus integrity and correction | went ahead 0 of 5 | went ahead 0 of 5 | went ahead 0 of 5 | went ahead 0 of 5 |
| Traffic drop triage | went ahead 0 of 5 | went ahead 0 of 5 | went ahead 0 of 5 | went ahead 0 of 5 |
No version pair on this site links these models. Every figure here was tested inside Claude Code.
How this comparison was made
Two models are compared on a job only where they answered exactly the same tasks, on the same test method, at the same level of difficulty, the same number of times. A page like this one exists only where that holds on at least three jobs. 11 jobs qualified here.
- Code review, in claude code: read from expansion-cohort.
- On-page audit, in claude code: read from expansion-cohort.
- Skill authoring, in claude code: read from expansion-cohort.
- Spec writing, in claude code: read from expansion-cohort.
- Link graph and metadata parity audit, in claude code: read from workflow-full-run.
- Corpus integrity and correction, in claude code: read from workflow-full-run.
- Traffic drop triage, in claude code: read from workflow-full-run.
- Getting clean JSON back, in claude code: read from tip-tier2-run.
- Getting the length right, in claude code: read from tip-tier2-run.
- Questions about recent events, in claude code: read from launch.
- Better step-by-step answers, in claude code: read from launch.
The same figures as data: /compare/claude-sonnet-vs-opus.json.