Claude vs GPT, Tested on the Same Tasks

Claude and GPT models were tested side by side in one test method only, the API. That is the only place the two makers share work, so it is the only place they are compared here.

Job by job

A solid mark is the higher score. “Too close to call” means the gap is smaller than the scores would move on another try.

Via API

JobClaude Haiku 4.5GPT-5 mini
On-page audit tier 123.9% too close to call $0.00085 a task same answer twice: one pass, not measured time: not timed44.9% best $0.00029 a task same answer twice: one pass, not measured time: not timed
Skill authoring tier 161.8% too close to call $0.00618 a task same answer twice: one pass, not measured time: not timed73.2% best $0.00297 a task same answer twice: one pass, not measured time: not timed
Spec writing tier 199.5% best $0.00332 a task same answer twice: one pass, not measured time: not timed94.5% $0.00248 a task same answer twice: one pass, not measured time: not timed
Questions about recent events tier 185.6% best Nothing was billed on this run for the version being priced here. same answer twice: 28 of 30 time: not timed48.6% Nothing was billed on this run for the version being priced here. same answer twice: 6 of 30 time: 1.5 s a call, across every tip asked five times
Better step-by-step answers tier 130.0% best Nothing was billed on this run for the version being priced here. same answer twice: 20 of 20 time: not timed16.0% too close to call Nothing was billed on this run for the version being priced here. same answer twice: 10 of 10 time: 1.5 s a call, across every tip asked five times

Where they differ

  • On spec writing, tested through the API, Claude Haiku 4.5 (99.5%) scored higher than GPT-5 mini (94.5%).
  • On questions about recent events, tested through the API, Claude Haiku 4.5 (85.6%) scored higher than GPT-5 mini (48.6%).

No version pair on this site links these models. Every figure here was tested through the API.

What is not compared, and why

Test methodClaude modelsGPT models
Via APIClaude Haiku 4.5GPT-5 mini, GPT-5.4 mini
In Claude CodeClaude Haiku 4.5, Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 5none
In CodexnoneGPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Terra, GPT-6 Luna, GPT-6 Sol

The Claude models in Claude Code and the GPT models in Codex are not comparable. Each was tested with different tools around it, and no model was tested in both. A score from one is never set against a score from the other.

The one link between test methods

A few models were tested two ways. The gap between their two scores on the same job shows how far apart the test methods sit. It is not used to move anyone’s score.

ModelJobFrom, toGapLikely range
Claude Haiku 4.5Accessibility auditVia API to In Claude Code First run, with and without the skill against Second run, every task twice+5.5 percentage points−11.7 to +22.8
Claude Haiku 4.5Accessibility auditVia API to In Claude Code First run, with and without the skill against Wider run, three ways+6.1 percentage points−12.1 to +24.3
Claude Haiku 4.5On-page auditVia API to In Claude Code First run, with and without the skill against Second run, every task twice−0.8 percentage points−28.4 to +26.8
Claude Haiku 4.5On-page auditVia API to In Claude Code First run, with and without the skill against Wider run, three ways−0.8 percentage points−30.3 to +28.7
Claude Haiku 4.5Skill authoringVia API to In Claude Code First run, with and without the skill against Second run, every task twice+8.2 percentage points−3.4 to +19.7
Claude Haiku 4.5Skill authoringVia API to In Claude Code First run, with and without the skill against Wider run, three ways+5.5 percentage points−6.7 to +17.6
Claude Haiku 4.5Spec writingVia API to In Claude Code First run, with and without the skill against Second run, every task twice−0.9 percentage points−2.2 to +0.4
Claude Haiku 4.5Spec writingVia API to In Claude Code First run, with and without the skill against Wider run, three ways−0.3 percentage points−1.3 to +0.7
Claude Haiku 4.5Getting clean JSON backVia API to In Claude Code Main run, tip by tip against Main run+0.0 percentage points+0.0 to +0.0
Claude Haiku 4.5Instructions before or afterVia API to In Claude Code Main run, tip by tip against Main run+0.0 percentage points+0.0 to +0.0
Claude Haiku 4.5Using tags and formattingVia API to In Claude Code Main run, tip by tip against Main run+0.0 percentage points+0.0 to +0.0
Claude Haiku 4.5Showing examplesVia API to In Claude Code Main run, tip by tip against Main run+0.0 percentage points−6.0 to +6.0
Claude Haiku 4.5Showing examplesVia API to In Claude Code Main run, tip by tip against Main run+0.0 percentage points−6.0 to +6.0
Claude Haiku 4.5Better step-by-step answersVia API to In Claude Code Main run, tip by tip against Main run−10.0 percentage points−27.0 to +7.0
Claude Haiku 4.5Stop made-up answersVia API to In Claude Code Main run, tip by tip against Main run+22.0 percentage points+10.4 to +33.6
Claude Haiku 4.5Instructions in long promptsVia API to In Claude Code Main run, tip by tip against Main run+0.0 percentage points+0.0 to +0.0
Claude Haiku 4.5Offering the model moneyVia API to In Claude Code Main run, tip by tip against Main run+37 words+3019.2 to +4440.8
Claude Haiku 4.5Questions about recent eventsVia API to In Claude Code Main run, tip by tip against Main run+11.1 percentage points−2.9 to +25.1
Claude Haiku 4.5Getting the length rightVia API to In Claude Code Main run, tip by tip against Main run+2.8 percentage points−33.3 to +38.9
GPT-5.4 miniQuestions about recent eventsVia API to In Codex API twin for the Codex run against Codex run, replaying two earlier tests−12.5 percentage pointsno range published
GPT-5.4 miniGetting the length rightVia API to In Codex API twin for the Codex run against Codex run, replaying two earlier tests+2.9 percentage pointsno range published
GPT-5.4 miniGetting the length rightVia API to In Codex API twin for the Codex run against Codex run, tip by tip+4.9 percentage pointsno range published
How this comparison was made

Two models are compared on a job only where they answered exactly the same tasks, on the same test method, at the same level of difficulty, the same number of times. A page like this one exists only where that holds on at least three jobs. Five jobs qualified here.

  • On-page audit, via api: read from skill-cohort.
  • Skill authoring, via api: read from skill-cohort.
  • Spec writing, via api: read from skill-cohort.
  • Questions about recent events, via api: read from launch.
  • Better step-by-step answers, via api: read from launch.

The same figures as data: /compare/claude-vs-gpt.json.