Claude vs GPT, Tested on the Same Tasks
Claude and GPT models were tested side by side in one test method only, the API. That is the only place the two makers share work, so it is the only place they are compared here.
Job by job
A solid mark is the higher score. “Too close to call” means the gap is smaller than the scores would move on another try.
Via API
| Job | Claude Haiku 4.5 | GPT-5 mini |
|---|---|---|
| On-page audit tier 1 | 23.9% too close to call $0.00085 a task same answer twice: one pass, not measured time: not timed | 44.9% best $0.00029 a task same answer twice: one pass, not measured time: not timed |
| Skill authoring tier 1 | 61.8% too close to call $0.00618 a task same answer twice: one pass, not measured time: not timed | 73.2% best $0.00297 a task same answer twice: one pass, not measured time: not timed |
| Spec writing tier 1 | 99.5% best $0.00332 a task same answer twice: one pass, not measured time: not timed | 94.5% $0.00248 a task same answer twice: one pass, not measured time: not timed |
| Questions about recent events tier 1 | 85.6% best Nothing was billed on this run for the version being priced here. same answer twice: 28 of 30 time: not timed | 48.6% Nothing was billed on this run for the version being priced here. same answer twice: 6 of 30 time: 1.5 s a call, across every tip asked five times |
| Better step-by-step answers tier 1 | 30.0% best Nothing was billed on this run for the version being priced here. same answer twice: 20 of 20 time: not timed | 16.0% too close to call Nothing was billed on this run for the version being priced here. same answer twice: 10 of 10 time: 1.5 s a call, across every tip asked five times |
Where they differ
- On spec writing, tested through the API, Claude Haiku 4.5 (99.5%) scored higher than GPT-5 mini (94.5%).
- On questions about recent events, tested through the API, Claude Haiku 4.5 (85.6%) scored higher than GPT-5 mini (48.6%).
No version pair on this site links these models. Every figure here was tested through the API.
What is not compared, and why
| Test method | Claude models | GPT models |
|---|---|---|
| Via API | Claude Haiku 4.5 | GPT-5 mini, GPT-5.4 mini |
| In Claude Code | Claude Haiku 4.5, Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 5 | none |
| In Codex | none | GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Terra, GPT-6 Luna, GPT-6 Sol |
The Claude models in Claude Code and the GPT models in Codex are not comparable. Each was tested with different tools around it, and no model was tested in both. A score from one is never set against a score from the other.
The one link between test methods
A few models were tested two ways. The gap between their two scores on the same job shows how far apart the test methods sit. It is not used to move anyone’s score.
| Model | Job | From, to | Gap | Likely range |
|---|---|---|---|---|
| Claude Haiku 4.5 | Accessibility audit | Via API to In Claude Code First run, with and without the skill against Second run, every task twice | +5.5 percentage points | −11.7 to +22.8 |
| Claude Haiku 4.5 | Accessibility audit | Via API to In Claude Code First run, with and without the skill against Wider run, three ways | +6.1 percentage points | −12.1 to +24.3 |
| Claude Haiku 4.5 | On-page audit | Via API to In Claude Code First run, with and without the skill against Second run, every task twice | −0.8 percentage points | −28.4 to +26.8 |
| Claude Haiku 4.5 | On-page audit | Via API to In Claude Code First run, with and without the skill against Wider run, three ways | −0.8 percentage points | −30.3 to +28.7 |
| Claude Haiku 4.5 | Skill authoring | Via API to In Claude Code First run, with and without the skill against Second run, every task twice | +8.2 percentage points | −3.4 to +19.7 |
| Claude Haiku 4.5 | Skill authoring | Via API to In Claude Code First run, with and without the skill against Wider run, three ways | +5.5 percentage points | −6.7 to +17.6 |
| Claude Haiku 4.5 | Spec writing | Via API to In Claude Code First run, with and without the skill against Second run, every task twice | −0.9 percentage points | −2.2 to +0.4 |
| Claude Haiku 4.5 | Spec writing | Via API to In Claude Code First run, with and without the skill against Wider run, three ways | −0.3 percentage points | −1.3 to +0.7 |
| Claude Haiku 4.5 | Getting clean JSON back | Via API to In Claude Code Main run, tip by tip against Main run | +0.0 percentage points | +0.0 to +0.0 |
| Claude Haiku 4.5 | Instructions before or after | Via API to In Claude Code Main run, tip by tip against Main run | +0.0 percentage points | +0.0 to +0.0 |
| Claude Haiku 4.5 | Using tags and formatting | Via API to In Claude Code Main run, tip by tip against Main run | +0.0 percentage points | +0.0 to +0.0 |
| Claude Haiku 4.5 | Showing examples | Via API to In Claude Code Main run, tip by tip against Main run | +0.0 percentage points | −6.0 to +6.0 |
| Claude Haiku 4.5 | Showing examples | Via API to In Claude Code Main run, tip by tip against Main run | +0.0 percentage points | −6.0 to +6.0 |
| Claude Haiku 4.5 | Better step-by-step answers | Via API to In Claude Code Main run, tip by tip against Main run | −10.0 percentage points | −27.0 to +7.0 |
| Claude Haiku 4.5 | Stop made-up answers | Via API to In Claude Code Main run, tip by tip against Main run | +22.0 percentage points | +10.4 to +33.6 |
| Claude Haiku 4.5 | Instructions in long prompts | Via API to In Claude Code Main run, tip by tip against Main run | +0.0 percentage points | +0.0 to +0.0 |
| Claude Haiku 4.5 | Offering the model money | Via API to In Claude Code Main run, tip by tip against Main run | +37 words | +3019.2 to +4440.8 |
| Claude Haiku 4.5 | Questions about recent events | Via API to In Claude Code Main run, tip by tip against Main run | +11.1 percentage points | −2.9 to +25.1 |
| Claude Haiku 4.5 | Getting the length right | Via API to In Claude Code Main run, tip by tip against Main run | +2.8 percentage points | −33.3 to +38.9 |
| GPT-5.4 mini | Questions about recent events | Via API to In Codex API twin for the Codex run against Codex run, replaying two earlier tests | −12.5 percentage points | no range published |
| GPT-5.4 mini | Getting the length right | Via API to In Codex API twin for the Codex run against Codex run, replaying two earlier tests | +2.9 percentage points | no range published |
| GPT-5.4 mini | Getting the length right | Via API to In Codex API twin for the Codex run against Codex run, tip by tip | +4.9 percentage points | no range published |
How this comparison was made
Two models are compared on a job only where they answered exactly the same tasks, on the same test method, at the same level of difficulty, the same number of times. A page like this one exists only where that holds on at least three jobs. Five jobs qualified here.
- On-page audit, via api: read from skill-cohort.
- Skill authoring, via api: read from skill-cohort.
- Spec writing, via api: read from skill-cohort.
- Questions about recent events, via api: read from launch.
- Better step-by-step answers, via api: read from launch.
The same figures as data: /compare/claude-vs-gpt.json.