{
  "surface": "openaddict-comparison",
  "version": "comparison-v1",
  "url": "https://openaddict.com/compare/claude-sonnet-vs-opus",
  "generatedAt": "2026-10-03",
  "generatedAtMeans": "The newest run date this document reads. It is not the time the file was written.",
  "rule": "Models are compared on a job only where the jobs-by-models grid holds a figure for each of them on one instrument at one tier, read from one run or from runs that asked exactly the same units. A model page exists only where that holds on at least 3 jobs. Who is ahead is the overlap rule over the compared models alone. Nothing crosses instruments.",
  "line": "11 jobs with the same work, tested inside Claude Code. They could not be told apart on 5 of the 7 that could be ranked.",
  "page": {
    "kind": "models",
    "copy": {
      "slug": "claude-sonnet-vs-opus",
      "kind": "models",
      "models": [
        "claude-sonnet-5",
        "claude-opus-5"
      ]
    },
    "models": [
      {
        "model": "claude-sonnet-5",
        "name": "Claude Sonnet 5"
      },
      {
        "model": "claude-opus-5",
        "name": "Claude Opus 5"
      }
    ],
    "shared": [
      {
        "job": {
          "id": "code-review",
          "kind": "skill-class",
          "label": "Code review",
          "href": "/models/compare#ts-skill-class-code-review",
          "scale": "unit",
          "measure": "share of the seeded code defects found"
        },
        "jobPath": "/jobs/code-review",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.7493749999999999,
            "shown": "74.9%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": "close",
            "cost": "$0.00406",
            "costUsd": 0.004063550000000001,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.9349999999999999,
            "shown": "93.5%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": "best",
            "cost": "$0.00948",
            "costUsd": 0.009484500000000002,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          }
        ],
        "state": "tied",
        "tied": true,
        "sentence": "On code review, tested inside Claude Code, they could not be told apart (Claude Sonnet 5 74.9%, Claude Opus 5 93.5%)."
      },
      {
        "job": {
          "id": "on-page-audit",
          "kind": "skill-class",
          "label": "On-page audit",
          "href": "/models/compare#ts-skill-class-on-page-audit",
          "scale": "unit",
          "measure": "share of the seeded on-page issues found"
        },
        "jobPath": "/jobs/on-page-audit",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.5297619047619048,
            "shown": "53.0%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": null,
            "cost": "$0.00362",
            "costUsd": 0.0036230000000000004,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.8805555555555555,
            "shown": "88.1%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": "best",
            "cost": "$0.00715",
            "costUsd": 0.0071491666666666665,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          }
        ],
        "state": "apart",
        "tied": false,
        "sentence": "On on-page audit, tested inside Claude Code, Claude Opus 5 (88.1%) scored higher than Claude Sonnet 5 (53.0%)."
      },
      {
        "job": {
          "id": "skill-authoring",
          "kind": "skill-class",
          "label": "Skill authoring",
          "href": "/models/compare#ts-skill-class-skill-authoring",
          "scale": "unit",
          "measure": "share of the specification requirements found"
        },
        "jobPath": "/jobs/skill-authoring",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 1,
            "shown": "100.0%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": "best",
            "cost": "$0.02027",
            "costUsd": 0.020272199999999997,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 1,
            "shown": "100.0%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": "best",
            "cost": "$0.05962",
            "costUsd": 0.05961924999999999,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          }
        ],
        "state": "tied",
        "tied": true,
        "sentence": "On skill authoring, tested inside Claude Code, they could not be told apart (Claude Sonnet 5 100.0%, Claude Opus 5 100.0%)."
      },
      {
        "job": {
          "id": "spec-writing",
          "kind": "skill-class",
          "label": "Spec writing",
          "href": "/models/compare#ts-skill-class-spec-writing",
          "scale": "unit",
          "measure": "share of the required sections and acceptance criteria found"
        },
        "jobPath": "/jobs/spec-writing",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.9621212121212122,
            "shown": "96.2%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": "best",
            "cost": "$0.01601",
            "costUsd": 0.016010433333333334,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.9227272727272726,
            "shown": "92.3%",
            "tier": "t1",
            "run": "expansion-cohort",
            "mark": null,
            "cost": "$0.06098",
            "costUsd": 0.06098191666666666,
            "repeat": "one pass, not measured",
            "repeatSetting": null,
            "time": "not timed"
          }
        ],
        "state": "apart",
        "tied": false,
        "sentence": "On spec writing, tested inside Claude Code, Claude Sonnet 5 (96.2%) scored higher than Claude Opus 5 (92.3%)."
      },
      {
        "job": {
          "id": "link-graph-and-metadata-parity-audit",
          "kind": "workflow",
          "label": "Link graph and metadata parity audit",
          "href": "/workflows/link-graph-and-metadata-parity-audit",
          "scale": "unit",
          "measure": "share of the planted faults found"
        },
        "jobPath": "/jobs/link-graph-audit",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.7372549019607844,
            "shown": "73.7%",
            "tier": "t1",
            "run": "workflow-full-run",
            "mark": null,
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "not asked twice",
            "repeatSetting": null,
            "time": "not timed"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.9294117647058824,
            "shown": "92.9%",
            "tier": "t1",
            "run": "workflow-full-run",
            "mark": "best",
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "not asked twice",
            "repeatSetting": null,
            "time": "not timed"
          }
        ],
        "state": "no-range",
        "tied": false,
        "sentence": "On link graph and metadata parity audit, tested inside Claude Code, the scores were Claude Sonnet 5 73.7% and Claude Opus 5 92.9%. No range was published, so whether that gap is real is not tested."
      },
      {
        "job": {
          "id": "corpus-integrity-and-correction",
          "kind": "workflow",
          "label": "Corpus integrity and correction",
          "href": "/workflows/corpus-integrity-and-correction",
          "scale": "unit",
          "measure": "share of the planted faults found"
        },
        "jobPath": "/jobs/corpus-integrity",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.6416666666666667,
            "shown": "64.2%",
            "tier": "t1",
            "run": "workflow-full-run",
            "mark": null,
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "not asked twice",
            "repeatSetting": null,
            "time": "not timed"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.7583333333333333,
            "shown": "75.8%",
            "tier": "t1",
            "run": "workflow-full-run",
            "mark": "best",
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "not asked twice",
            "repeatSetting": null,
            "time": "not timed"
          }
        ],
        "state": "no-range",
        "tied": false,
        "sentence": "On corpus integrity and correction, tested inside Claude Code, the scores were Claude Sonnet 5 64.2% and Claude Opus 5 75.8%. No range was published, so whether that gap is real is not tested."
      },
      {
        "job": {
          "id": "traffic-drop-triage",
          "kind": "workflow",
          "label": "Traffic drop triage",
          "href": "/workflows/traffic-drop-triage",
          "scale": "unit",
          "measure": "share of runs that named the right cause"
        },
        "jobPath": "/jobs/traffic-drop-triage",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.8,
            "shown": "80.0%",
            "tier": "t1",
            "run": "workflow-full-run",
            "mark": null,
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "not asked twice",
            "repeatSetting": null,
            "time": "not timed"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 1,
            "shown": "100.0%",
            "tier": "t1",
            "run": "workflow-full-run",
            "mark": "best",
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "not asked twice",
            "repeatSetting": null,
            "time": "not timed"
          }
        ],
        "state": "no-range",
        "tied": false,
        "sentence": "On traffic drop triage, tested inside Claude Code, the scores were Claude Sonnet 5 80.0% and Claude Opus 5 100.0%. No range was published, so whether that gap is real is not tested."
      },
      {
        "job": {
          "id": "C01-json-schema",
          "kind": "tip",
          "label": "Getting clean JSON back",
          "href": "/models/compare#ts-tip-single-turn-C01-json-schema",
          "scale": "unit",
          "measure": "score with no tip applied"
        },
        "jobPath": "/jobs/json-extraction",
        "instrument": "claude-code",
        "tier": "t2",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0,
            "shown": "0.0%",
            "tier": "t2",
            "run": "tip-tier2-run",
            "mark": null,
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "10 of 10",
            "repeatSetting": "no sampling control here",
            "time": "11.8 s a call, across every tip asked five times"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0,
            "shown": "0.0%",
            "tier": "t2",
            "run": "tip-tier2-run",
            "mark": null,
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "10 of 10",
            "repeatSetting": "no sampling control here",
            "time": "12.0 s a call, across every tip asked five times"
          }
        ],
        "state": "none",
        "tied": true,
        "sentence": "On getting clean json back, tested inside Claude Code, none of them managed it with no help."
      },
      {
        "job": {
          "id": "C17-exact-length",
          "kind": "tip",
          "label": "Getting the length right",
          "href": "/models/compare#ts-tip-single-turn-C17-exact-length",
          "scale": "unit",
          "measure": "score with no tip applied"
        },
        "jobPath": "/jobs/length-control",
        "instrument": "claude-code",
        "tier": "t2",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.014166666666666671,
            "shown": "1.4%",
            "tier": "t2",
            "run": "tip-tier2-run",
            "mark": "best",
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "7 of 10",
            "repeatSetting": "no sampling control here",
            "time": "11.8 s a call, across every tip asked five times"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.0025000000000000022,
            "shown": "0.3%",
            "tier": "t2",
            "run": "tip-tier2-run",
            "mark": "close",
            "cost": "no cost published",
            "costUsd": null,
            "repeat": "8 of 10",
            "repeatSetting": "no sampling control here",
            "time": "12.0 s a call, across every tip asked five times"
          }
        ],
        "state": "tied",
        "tied": true,
        "sentence": "On getting the length right, tested inside Claude Code, they could not be told apart (Claude Sonnet 5 1.4%, Claude Opus 5 0.3%)."
      },
      {
        "job": {
          "id": "C16-cutoff-disclosure",
          "kind": "tip",
          "label": "Questions about recent events",
          "href": "/models/compare#ts-tip-single-turn-C16-cutoff-disclosure",
          "scale": "unit",
          "measure": "score with no tip applied"
        },
        "jobPath": "/jobs/recent-events",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.9333333333333333,
            "shown": "93.3%",
            "tier": "t1",
            "run": "launch",
            "mark": "best",
            "cost": "$0.00757",
            "costUsd": 0.007570266666666667,
            "repeat": "22 of 30",
            "repeatSetting": "no sampling control here",
            "time": "11.8 s a call, across every tip asked five times"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.8027777777777778,
            "shown": "80.3%",
            "tier": "t1",
            "run": "launch",
            "mark": "close",
            "cost": "$0.01566",
            "costUsd": 0.015659,
            "repeat": "23 of 30",
            "repeatSetting": "no sampling control here",
            "time": "12.0 s a call, across every tip asked five times"
          }
        ],
        "state": "tied",
        "tied": true,
        "sentence": "On questions about recent events, tested inside Claude Code, they could not be told apart (Claude Sonnet 5 93.3%, Claude Opus 5 80.3%)."
      },
      {
        "job": {
          "id": "C05-think-step-by-step",
          "kind": "tip",
          "label": "Better step-by-step answers",
          "href": "/models/compare#ts-tip-single-turn-C05-think-step-by-step",
          "scale": "unit",
          "measure": "score with no tip applied"
        },
        "jobPath": "/jobs/step-by-step",
        "instrument": "claude-code",
        "tier": "t1",
        "figures": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "mean": 0.4,
            "shown": "40.0%",
            "tier": "t1",
            "run": "launch",
            "mark": "best",
            "cost": "$0.00630",
            "costUsd": 0.006297,
            "repeat": "10 of 10",
            "repeatSetting": "no sampling control here",
            "time": "11.8 s a call, across every tip asked five times"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "mean": 0.4,
            "shown": "40.0%",
            "tier": "t1",
            "run": "launch",
            "mark": "best",
            "cost": "$0.00807",
            "costUsd": 0.00807,
            "repeat": "10 of 10",
            "repeatSetting": "no sampling control here",
            "time": "12.0 s a call, across every tip asked five times"
          }
        ],
        "state": "tied",
        "tied": true,
        "sentence": "On better step-by-step answers, tested inside Claude Code, they could not be told apart (Claude Sonnet 5 40.0%, Claude Opus 5 40.0%)."
      }
    ],
    "jobs": 11,
    "tiedCount": 5,
    "rankable": 7,
    "differ": [
      "On on-page audit, tested inside Claude Code, Claude Opus 5 (88.1%) scored higher than Claude Sonnet 5 (53.0%).",
      "On spec writing, tested inside Claude Code, Claude Sonnet 5 (96.2%) scored higher than Claude Opus 5 (92.3%)."
    ],
    "untested": [
      "On link graph and metadata parity audit, tested inside Claude Code, the scores were Claude Sonnet 5 73.7% and Claude Opus 5 92.9%. No range was published, so whether that gap is real is not tested.",
      "On corpus integrity and correction, tested inside Claude Code, the scores were Claude Sonnet 5 64.2% and Claude Opus 5 75.8%. No range was published, so whether that gap is real is not tested.",
      "On traffic drop triage, tested inside Claude Code, the scores were Claude Sonnet 5 80.0% and Claude Opus 5 100.0%. No range was published, so whether that gap is real is not tested."
    ],
    "stops": [
      {
        "workflow": "link-graph-and-metadata-parity-audit",
        "workflowName": "Link graph and metadata parity audit",
        "rows": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "notTold": "went ahead 5 of 5",
            "told": "went ahead 0 of 5"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "notTold": "went ahead 0 of 5",
            "told": "went ahead 0 of 5"
          }
        ]
      },
      {
        "workflow": "corpus-integrity-and-correction",
        "workflowName": "Corpus integrity and correction",
        "rows": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "notTold": "went ahead 0 of 5",
            "told": "went ahead 0 of 5"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "notTold": "went ahead 0 of 5",
            "told": "went ahead 0 of 5"
          }
        ]
      },
      {
        "workflow": "traffic-drop-triage",
        "workflowName": "Traffic drop triage",
        "rows": [
          {
            "model": "claude-sonnet-5",
            "name": "Claude Sonnet 5",
            "notTold": "went ahead 0 of 5",
            "told": "went ahead 0 of 5"
          },
          {
            "model": "claude-opus-5",
            "name": "Claude Opus 5",
            "notTold": "went ahead 0 of 5",
            "told": "went ahead 0 of 5"
          }
        ]
      }
    ],
    "version": null,
    "methods": [
      "claude-code"
    ]
  }
}
