Claude Sonnet vs Opus, Tested on the Same Tasks

Claude Sonnet 5 and Claude Opus 5 did the same work on 11 jobs. On 7 of them the scores could be ranked, and on 5 of those they could not be told apart.

Job by job

A solid mark is the higher score. “Too close to call” means the gap is smaller than the scores would move on another try.

In Claude Code

JobClaude Sonnet 5Claude Opus 5
Code review tier 174.9% too close to call $0.00406 a task same answer twice: one pass, not measured time: not timed93.5% best $0.00948 a task same answer twice: one pass, not measured time: not timed
On-page audit tier 153.0% $0.00362 a task same answer twice: one pass, not measured time: not timed88.1% best $0.00715 a task same answer twice: one pass, not measured time: not timed
Skill authoring tier 1100.0% best $0.02027 a task same answer twice: one pass, not measured time: not timed100.0% best $0.05962 a task same answer twice: one pass, not measured time: not timed
Spec writing tier 196.2% best $0.01601 a task same answer twice: one pass, not measured time: not timed92.3% $0.06098 a task same answer twice: one pass, not measured time: not timed
Link graph and metadata parity audit tier 173.7% no cost published same answer twice: not asked twice time: not timed92.9% best no cost published same answer twice: not asked twice time: not timed
Corpus integrity and correction tier 164.2% no cost published same answer twice: not asked twice time: not timed75.8% best no cost published same answer twice: not asked twice time: not timed
Traffic drop triage tier 180.0% no cost published same answer twice: not asked twice time: not timed100.0% best no cost published same answer twice: not asked twice time: not timed
Getting clean JSON back tier 20.0% no cost published same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times0.0% no cost published same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times
Getting the length right tier 21.4% best no cost published same answer twice: 7 of 10 time: 11.8 s a call, across every tip asked five times0.3% too close to call no cost published same answer twice: 8 of 10 time: 12.0 s a call, across every tip asked five times
Questions about recent events tier 193.3% best $0.00757 a task same answer twice: 22 of 30 time: 11.8 s a call, across every tip asked five times80.3% too close to call $0.01566 a task same answer twice: 23 of 30 time: 12.0 s a call, across every tip asked five times
Better step-by-step answers tier 140.0% best $0.00630 a task same answer twice: 10 of 10 time: 11.8 s a call, across every tip asked five times40.0% best $0.00807 a task same answer twice: 10 of 10 time: 12.0 s a call, across every tip asked five times

Where they differ

  • On on-page audit, tested inside Claude Code, Claude Opus 5 (88.1%) scored higher than Claude Sonnet 5 (53.0%).
  • On spec writing, tested inside Claude Code, Claude Sonnet 5 (96.2%) scored higher than Claude Opus 5 (92.3%).

Gaps with no range

  • On link graph and metadata parity audit, tested inside Claude Code, the scores were Claude Sonnet 5 73.7% and Claude Opus 5 92.9%. No range was published, so whether that gap is real is not tested.
  • On corpus integrity and correction, tested inside Claude Code, the scores were Claude Sonnet 5 64.2% and Claude Opus 5 75.8%. No range was published, so whether that gap is real is not tested.
  • On traffic drop triage, tested inside Claude Code, the scores were Claude Sonnet 5 80.0% and Claude Opus 5 100.0%. No range was published, so whether that gap is real is not tested.

When an input was missing

One input was left out on purpose. In one run nothing said so. In the other, the model was told.

JobClaude Sonnet 5, not toldClaude Opus 5, not toldClaude Sonnet 5, toldClaude Opus 5, told
Link graph and metadata parity auditwent ahead 5 of 5went ahead 0 of 5went ahead 0 of 5went ahead 0 of 5
Corpus integrity and correctionwent ahead 0 of 5went ahead 0 of 5went ahead 0 of 5went ahead 0 of 5
Traffic drop triagewent ahead 0 of 5went ahead 0 of 5went ahead 0 of 5went ahead 0 of 5

No version pair on this site links these models. Every figure here was tested inside Claude Code.

How this comparison was made

Two models are compared on a job only where they answered exactly the same tasks, on the same test method, at the same level of difficulty, the same number of times. A page like this one exists only where that holds on at least three jobs. 11 jobs qualified here.

  • Code review, in claude code: read from expansion-cohort.
  • On-page audit, in claude code: read from expansion-cohort.
  • Skill authoring, in claude code: read from expansion-cohort.
  • Spec writing, in claude code: read from expansion-cohort.
  • Link graph and metadata parity audit, in claude code: read from workflow-full-run.
  • Corpus integrity and correction, in claude code: read from workflow-full-run.
  • Traffic drop triage, in claude code: read from workflow-full-run.
  • Getting clean JSON back, in claude code: read from tip-tier2-run.
  • Getting the length right, in claude code: read from tip-tier2-run.
  • Questions about recent events, in claude code: read from launch.
  • Better step-by-step answers, in claude code: read from launch.

The same figures as data: /compare/claude-sonnet-vs-opus.json.