Which AI Model to Use for Each Job

Four picks for each job. Each pick is made inside one test method, and the run behind it is named. The same picks are in a data file, as of 2026-09-19.

How to use these picks

  1. Choose a model for each job, not one model for every job.
  2. For each job, take the cheapest model that is too close to call against the best, unless it repeats itself less often or answers more slowly.
  3. Never compare scores from two test methods, except through a model tested both ways, like Claude Haiku 4.5.
  4. Fetch this page or its data file again before you rely on it, because the models on offer and their versions change.
  5. A tier means we made the tasks harder until the models could be told apart.
  6. “Not run” means we never tested that model on that job, not that it scored zero.

The picks, job by job

JobBest on its ownCheapest within range of the bestMost consistentFastest
Accessibility audit Scored as the share of the seeded accessibility defects found In the API, each model was given tasks at its own level of difficulty, so no one model is ranked first. In Claude Code, each model was given tasks at its own level of difficulty, so no one model is ranked first. Those test methods are never compared with each other.
Via API Each model was given tasks at its own level of difficulty, so no one model is ranked first.
In Claude Code Each model was given tasks at its own level of difficulty, so no one model is ranked first.
Via API Each model was given tasks at its own level of difficulty, so no one model is ranked first.
In Claude Code Each model was given tasks at its own level of difficulty, so no one model is ranked first.
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code Claude Opus 5, same answer on 38 of 40 no sampling control here, Second run, every task twice
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.
Code review Scored as the share of the seeded code defects found In the API, GPT-5 mini scores highest. Gemini 3.1 Flash Lite is too close to call. In Claude Code, Claude Opus 5 scores highest. Claude Fable 5.1 and Claude Sonnet 5 are too close to call. Those test methods are never compared with each other.
Via API GPT-5 mini 70.0% Too close to call with Gemini 3.1 Flash Lite tier 1, Wider run, three ways
In Claude Code Claude Opus 5 93.5% Too close to call with Claude Fable 5.1 and Claude Sonnet 5 tier 1, Wider run, three ways
Via API The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
In Claude Code The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code No run asked these tasks more than once this way, so there is no repeat to compare.
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.
On-page audit Scored as the share of the seeded on-page issues found In the API, Gemini 3.1 Flash Lite scores highest. GPT-5 mini and Claude Haiku 4.5 are too close to call. In Claude Code, Claude Fable 5.1 scores highest. Claude Opus 5 is too close to call. Those test methods are never compared with each other.
Via API Gemini 3.1 Flash Lite 48.8% Too close to call with GPT-5 mini and Claude Haiku 4.5 tier 1, First run, with and without the skill
In Claude Code Claude Fable 5.1 97.2% Too close to call with Claude Opus 5 tier 1, Wider run, three ways
Via API Gemini 3.1 Flash Lite $0.00023 a task The best score is also the cheapest in range billed through the API
In Claude Code The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code Claude Opus 5, same answer on 33 of 40 no sampling control here, Second run, every task twice
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.
Skill authoring Scored as the share of the specification requirements found In the API, Gemini 3.1 Flash Lite scores highest. GPT-5 mini is too close to call. In Claude Code, Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5 tie for the top score. Those test methods are never compared with each other.
Via API Gemini 3.1 Flash Lite 87.3% Too close to call with GPT-5 mini tier 1, First run, with and without the skill
In Claude Code Claude Fable 5.1 100.0%, Claude Opus 5 100.0%, Claude Sonnet 5 100.0%, tied tier 1, Wider run, three ways
Via API Gemini 3.1 Flash Lite $0.00088 a task The best score is also the cheapest in range billed through the API
In Claude Code The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code Claude Opus 5, same answer on 28 of 29 no sampling control here, Second run, every task twice
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.
Spec writing Scored as the share of the required sections and acceptance criteria found In the API, Claude Haiku 4.5 scores highest. Gemini 3.1 Flash Lite is too close to call. Gemini 3.1 Flash Lite is the cheapest of those. In Claude Code, Claude Haiku 4.5 scores highest. Claude Fable 5.1 is too close to call. Those test methods are never compared with each other.
Via API Claude Haiku 4.5 99.5% Too close to call with Gemini 3.1 Flash Lite tier 1, First run, with and without the skill
In Claude Code Claude Haiku 4.5 99.2% Too close to call with Claude Fable 5.1 tier 1, Wider run, three ways
Via API Gemini 3.1 Flash Lite $0.00080 a task Against Claude Haiku 4.5 $0.00332 a task, the best score billed through the API
In Claude Code The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code Claude Opus 5, same answer on 27 of 40 no sampling control here, Second run, every task twice
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.
Link graph and metadata parity audit Scored as the share of the planted faults found In Claude Code, Claude Fable 5.1 scores highest.
In Claude Code Claude Fable 5.1 100.0% No range published, so nobody is placed against it tier 1, Workflow run, five times each
In Claude Code This run does not publish a cost per task.
In Claude Code No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code No run of this job timed its calls this way.
Corpus integrity and correction Scored as the share of the planted faults found In Claude Code, Claude Opus 5 scores highest.
In Claude Code Claude Opus 5 75.8% No range published, so nobody is placed against it tier 1, Workflow run, five times each
In Claude Code This run does not publish a cost per task.
In Claude Code No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code No run of this job timed its calls this way.
Traffic drop triage Scored as the share of runs that named the right cause In Claude Code, Claude Opus 5 and Claude Fable 5.1 tie for the top score.
In Claude Code Claude Opus 5 100.0%, Claude Fable 5.1 100.0%, tied No range published, so nobody is placed against it tier 1, Workflow run, five times each
In Claude Code This run does not publish a cost per task.
In Claude Code No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code No run of this job timed its calls this way.
Getting clean JSON back Scored as the score with no tip applied In the API, no model managed this on its own. In Claude Code, no model managed this on its own. In Codex, no model managed this on its own. Those test methods are never compared with each other.
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.
Getting the length right Scored as the score with no tip applied In the API, GPT-5 mini scores highest. Gemini 3.1 Flash Lite is too close to call. In Claude Code, Claude Haiku 4.5 scores highest. Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1 are too close to call. In Codex, GPT-5.6 Terra scores highest. Those test methods are never compared with each other.
Via API GPT-5 mini 5.0% Too close to call with Gemini 3.1 Flash Lite tier 2, Harder tasks, tier 2
In Claude Code Claude Haiku 4.5 4.2% Too close to call with Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1 tier 2, Harder tasks, tier 2
In Codex GPT-5.6 Terra 62.4% No range published, so nobody is placed against it tier 1, Codex run, tip by tip
Via API This run does not publish a cost per task.
In Claude Code This run does not publish a cost per task.
In Codex The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked.
In Claude Code Claude Opus 5, same answer on 8 of 10 Too close to call with Claude Haiku 4.5, Claude Fable 5.1 and Claude Sonnet 5 no sampling control here, Each tip asked five times
In Codex No run asked these tasks more than once this way, so there is no repeat to compare.
Via API Gemini 3.1 Flash Lite 563 ms Median time per call, across all twelve tips asked five times 2 models timed, Each tip asked five times
In Claude Code Claude Haiku 4.5 11.4 s Median time per call, across all twelve tips asked five times 4 models timed, Each tip asked five times
In Codex No run of this job timed its calls this way.
Questions about recent events Scored as the score with no tip applied In the API, Claude Haiku 4.5 scores highest. In Claude Code, Claude Haiku 4.5 scores highest. Claude Sonnet 5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 are too close to call. In Codex, only one model was measured on this run, so there is nothing to rank. Those test methods are never compared with each other.
Via API Claude Haiku 4.5 85.6% tier 1, Main run, tip by tip
In Claude Code Claude Haiku 4.5 96.7% Too close to call with Claude Sonnet 5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 tier 1, Main run
In Codex Only one model was measured on this run, so there is nothing to rank.
Via API The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
In Claude Code Claude Haiku 4.5 $0.00220 a task The best score is also the cheapest in range run on subscription; shown at API list price for comparison
In Codex Only one model was measured on this run, so there is nothing to rank.
Via API The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked.
In Claude Code Claude Haiku 4.5, same answer on 28 of 30 Too close to call with Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5 no sampling control here, Each tip asked five times
In Codex No run asked these tasks more than once this way, so there is no repeat to compare.
Via API Gemini 3.1 Flash Lite 563 ms Median time per call, across all twelve tips asked five times 2 models timed, Each tip asked five times
In Claude Code Claude Haiku 4.5 11.4 s Median time per call, across all twelve tips asked five times 4 models timed, Each tip asked five times
In Codex No run of this job timed its calls this way.
Better step-by-step answers Scored as the score with no tip applied In the API, Gemini 3.1 Flash Lite scores highest. Claude Haiku 4.5 is too close to call. In Claude Code, Claude Fable 5, Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5 tie for the top score. Claude Haiku 4.5 is too close to call. Claude Haiku 4.5 is the cheapest of those. In Codex, GPT-5.4 mini, GPT-5.6 Luna and GPT-5.6 Terra tie for the top score. Those test methods are never compared with each other.
Via API Gemini 3.1 Flash Lite 40.0% Too close to call with Claude Haiku 4.5 tier 1, Main run, tip by tip
In Claude Code Claude Fable 5 40.0%, Claude Fable 5.1 40.0%, Claude Opus 5 40.0%, Claude Sonnet 5 40.0%, tied Too close to call with Claude Haiku 4.5 tier 1, Main run
In Codex GPT-5.4 mini 30.0%, GPT-5.6 Luna 30.0%, GPT-5.6 Terra 30.0%, tied No range published, so nobody is placed against it tier 1, Codex run, tip by tip
Via API The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
In Claude Code Claude Haiku 4.5 $0.00135 a task Against Claude Opus 5 $0.00807 a task, the best score run on subscription; shown at API list price for comparison
In Codex The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked.
In Claude Code Claude Fable 5.1, same answer on 10 of 10; Claude Opus 5, same answer on 10 of 10; Claude Sonnet 5, same answer on 10 of 10, tied Too close to call with Claude Haiku 4.5 no sampling control here, Each tip asked five times
In Codex No run asked these tasks more than once this way, so there is no repeat to compare.
Via API Gemini 3.1 Flash Lite 563 ms Median time per call, across all twelve tips asked five times 2 models timed, Each tip asked five times
In Claude Code Claude Haiku 4.5 11.4 s Median time per call, across all twelve tips asked five times 4 models timed, Each tip asked five times
In Codex No run of this job timed its calls this way.