Which AI Model to Use for Each Job
Four picks for each job. Each pick is made inside one test method, and the run behind it is named. The same picks are in a data file, as of 2026-09-19.
How to use these picks
- Choose a model for each job, not one model for every job.
- For each job, take the cheapest model that is too close to call against the best, unless it repeats itself less often or answers more slowly.
- Never compare scores from two test methods, except through a model tested both ways, like Claude Haiku 4.5.
- Fetch this page or its data file again before you rely on it, because the models on offer and their versions change.
- A tier means we made the tasks harder until the models could be told apart.
- “Not run” means we never tested that model on that job, not that it scored zero.
The picks, job by job
| Job | Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|---|
| Accessibility audit Scored as the share of the seeded accessibility defects found In the API, each model was given tasks at its own level of difficulty, so no one model is ranked first. In Claude Code, each model was given tasks at its own level of difficulty, so no one model is ranked first. Those test methods are never compared with each other. | Via API
Each model was given tasks at its own level of difficulty, so no one model is ranked first. In Claude Code
Each model was given tasks at its own level of difficulty, so no one model is ranked first. | Via API
Each model was given tasks at its own level of difficulty, so no one model is ranked first. In Claude Code
Each model was given tasks at its own level of difficulty, so no one model is ranked first. | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 38 of 40
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
| Code review Scored as the share of the seeded code defects found In the API, GPT-5 mini scores highest. Gemini 3.1 Flash Lite is too close to call. In Claude Code, Claude Opus 5 scores highest. Claude Fable 5.1 and Claude Sonnet 5 are too close to call. Those test methods are never compared with each other. | Via API
GPT-5 mini 70.0%
Too close to call with Gemini 3.1 Flash Lite
tier 1, Wider run, three ways In Claude Code
Claude Opus 5 93.5%
Too close to call with Claude Fable 5.1 and Claude Sonnet 5
tier 1, Wider run, three ways | Via API
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. In Claude Code
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
| On-page audit Scored as the share of the seeded on-page issues found In the API, Gemini 3.1 Flash Lite scores highest. GPT-5 mini and Claude Haiku 4.5 are too close to call. In Claude Code, Claude Fable 5.1 scores highest. Claude Opus 5 is too close to call. Those test methods are never compared with each other. | Via API
Gemini 3.1 Flash Lite 48.8%
Too close to call with GPT-5 mini and Claude Haiku 4.5
tier 1, First run, with and without the skill In Claude Code
Claude Fable 5.1 97.2%
Too close to call with Claude Opus 5
tier 1, Wider run, three ways | Via API
Gemini 3.1 Flash Lite $0.00023 a task
The best score is also the cheapest in range
billed through the API In Claude Code
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 33 of 40
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
| Skill authoring Scored as the share of the specification requirements found In the API, Gemini 3.1 Flash Lite scores highest. GPT-5 mini is too close to call. In Claude Code, Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5 tie for the top score. Those test methods are never compared with each other. | Via API
Gemini 3.1 Flash Lite 87.3%
Too close to call with GPT-5 mini
tier 1, First run, with and without the skill In Claude Code
Claude Fable 5.1 100.0%, Claude Opus 5 100.0%, Claude Sonnet 5 100.0%, tied
tier 1, Wider run, three ways | Via API
Gemini 3.1 Flash Lite $0.00088 a task
The best score is also the cheapest in range
billed through the API In Claude Code
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 28 of 29
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
| Spec writing Scored as the share of the required sections and acceptance criteria found In the API, Claude Haiku 4.5 scores highest. Gemini 3.1 Flash Lite is too close to call. Gemini 3.1 Flash Lite is the cheapest of those. In Claude Code, Claude Haiku 4.5 scores highest. Claude Fable 5.1 is too close to call. Those test methods are never compared with each other. | Via API
Claude Haiku 4.5 99.5%
Too close to call with Gemini 3.1 Flash Lite
tier 1, First run, with and without the skill In Claude Code
Claude Haiku 4.5 99.2%
Too close to call with Claude Fable 5.1
tier 1, Wider run, three ways | Via API
Gemini 3.1 Flash Lite $0.00080 a task
Against Claude Haiku 4.5 $0.00332 a task, the best score
billed through the API In Claude Code
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 27 of 40
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
| Link graph and metadata parity audit Scored as the share of the planted faults found In Claude Code, Claude Fable 5.1 scores highest. | In Claude Code
Claude Fable 5.1 100.0%
No range published, so nobody is placed against it
tier 1, Workflow run, five times each | In Claude Code
This run does not publish a cost per task. | In Claude Code
No run asked these tasks more than once this way, so there is no repeat to compare. | In Claude Code
No run of this job timed its calls this way. |
| Corpus integrity and correction Scored as the share of the planted faults found In Claude Code, Claude Opus 5 scores highest. | In Claude Code
Claude Opus 5 75.8%
No range published, so nobody is placed against it
tier 1, Workflow run, five times each | In Claude Code
This run does not publish a cost per task. | In Claude Code
No run asked these tasks more than once this way, so there is no repeat to compare. | In Claude Code
No run of this job timed its calls this way. |
| Traffic drop triage Scored as the share of runs that named the right cause In Claude Code, Claude Opus 5 and Claude Fable 5.1 tie for the top score. | In Claude Code
Claude Opus 5 100.0%, Claude Fable 5.1 100.0%, tied
No range published, so nobody is placed against it
tier 1, Workflow run, five times each | In Claude Code
This run does not publish a cost per task. | In Claude Code
No run asked these tasks more than once this way, so there is no repeat to compare. | In Claude Code
No run of this job timed its calls this way. |
| Getting clean JSON back Scored as the score with no tip applied In the API, no model managed this on its own. In Claude Code, no model managed this on its own. In Codex, no model managed this on its own. Those test methods are never compared with each other. | Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. | Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. | Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. | Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. |
| Getting the length right Scored as the score with no tip applied In the API, GPT-5 mini scores highest. Gemini 3.1 Flash Lite is too close to call. In Claude Code, Claude Haiku 4.5 scores highest. Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1 are too close to call. In Codex, GPT-5.6 Terra scores highest. Those test methods are never compared with each other. | Via API
GPT-5 mini 5.0%
Too close to call with Gemini 3.1 Flash Lite
tier 2, Harder tasks, tier 2 In Claude Code
Claude Haiku 4.5 4.2%
Too close to call with Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1
tier 2, Harder tasks, tier 2 In Codex
GPT-5.6 Terra 62.4%
No range published, so nobody is placed against it
tier 1, Codex run, tip by tip | Via API
This run does not publish a cost per task. In Claude Code
This run does not publish a cost per task. In Codex
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked. In Claude Code
Claude Opus 5, same answer on 8 of 10
Too close to call with Claude Haiku 4.5, Claude Fable 5.1 and Claude Sonnet 5
no sampling control here, Each tip asked five times In Codex
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
Gemini 3.1 Flash Lite 563 ms
Median time per call, across all twelve tips asked five times
2 models timed, Each tip asked five times In Claude Code
Claude Haiku 4.5 11.4 s
Median time per call, across all twelve tips asked five times
4 models timed, Each tip asked five times In Codex
No run of this job timed its calls this way. |
| Questions about recent events Scored as the score with no tip applied In the API, Claude Haiku 4.5 scores highest. In Claude Code, Claude Haiku 4.5 scores highest. Claude Sonnet 5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 are too close to call. In Codex, only one model was measured on this run, so there is nothing to rank. Those test methods are never compared with each other. | Via API
Claude Haiku 4.5 85.6%
tier 1, Main run, tip by tip In Claude Code
Claude Haiku 4.5 96.7%
Too close to call with Claude Sonnet 5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5
tier 1, Main run In Codex
Only one model was measured on this run, so there is nothing to rank. | Via API
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. In Claude Code
Claude Haiku 4.5 $0.00220 a task
The best score is also the cheapest in range
run on subscription; shown at API list price for comparison In Codex
Only one model was measured on this run, so there is nothing to rank. | Via API
The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked. In Claude Code
Claude Haiku 4.5, same answer on 28 of 30
Too close to call with Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5
no sampling control here, Each tip asked five times In Codex
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
Gemini 3.1 Flash Lite 563 ms
Median time per call, across all twelve tips asked five times
2 models timed, Each tip asked five times In Claude Code
Claude Haiku 4.5 11.4 s
Median time per call, across all twelve tips asked five times
4 models timed, Each tip asked five times In Codex
No run of this job timed its calls this way. |
| Better step-by-step answers Scored as the score with no tip applied In the API, Gemini 3.1 Flash Lite scores highest. Claude Haiku 4.5 is too close to call. In Claude Code, Claude Fable 5, Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5 tie for the top score. Claude Haiku 4.5 is too close to call. Claude Haiku 4.5 is the cheapest of those. In Codex, GPT-5.4 mini, GPT-5.6 Luna and GPT-5.6 Terra tie for the top score. Those test methods are never compared with each other. | Via API
Gemini 3.1 Flash Lite 40.0%
Too close to call with Claude Haiku 4.5
tier 1, Main run, tip by tip In Claude Code
Claude Fable 5 40.0%, Claude Fable 5.1 40.0%, Claude Opus 5 40.0%, Claude Sonnet 5 40.0%, tied
Too close to call with Claude Haiku 4.5
tier 1, Main run In Codex
GPT-5.4 mini 30.0%, GPT-5.6 Luna 30.0%, GPT-5.6 Terra 30.0%, tied
No range published, so nobody is placed against it
tier 1, Codex run, tip by tip | Via API
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. In Claude Code
Claude Haiku 4.5 $0.00135 a task
Against Claude Opus 5 $0.00807 a task, the best score
run on subscription; shown at API list price for comparison In Codex
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked. In Claude Code
Claude Fable 5.1, same answer on 10 of 10; Claude Opus 5, same answer on 10 of 10; Claude Sonnet 5, same answer on 10 of 10, tied
Too close to call with Claude Haiku 4.5
no sampling control here, Each tip asked five times In Codex
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
Gemini 3.1 Flash Lite 563 ms
Median time per call, across all twelve tips asked five times
2 models timed, Each tip asked five times In Claude Code
Claude Haiku 4.5 11.4 s
Median time per call, across all twelve tips asked five times
4 models timed, Each tip asked five times In Codex
No run of this job timed its calls this way. |