Best LLM for Code Review, Tested

A code review means reading a change to some code and finding the mistakes that were put in it. Each model was given the same ten tasks with no help, tested in the API and Claude Code.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API GPT-5 mini 70.0% Too close to call with Gemini 3.1 Flash Lite tier 1, Wider run, three ways
In Claude Code Claude Opus 5 93.5% Too close to call with Claude Fable 5.1, Claude Sonnet 5, Claude Opus 5.5 and Claude Fable 5 tier 1, Wider run, three ways Some of these models come from another run of the same tasks
Via API The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
In Claude Code Claude Sonnet 5 $0.00406 a task Against Claude Opus 5 $0.00948 a task, the best score run on subscription; shown at API list price for comparison
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code No run asked these tasks more than once this way, so there is no repeat to compare.
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.5not run
GPT-5 mini70.0% tier 1 bestone pass, not measured
Gemini 3.1 Flash Lite57.5% tier 1 too close to callone pass, not measured
GPT-5.4 mininot run

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.558.3% tier 1one pass, not measured
Claude Fable 585.8% tier 1 too close to callone pass, not measured
Claude Fable 5.189.0% tier 1 too close to callone pass, not measured
Claude Opus 593.5% tier 1 bestone pass, not measured
Claude Opus 5.585.4% tier 1 too close to callone pass, not measured
Claude Sonnet 574.9% tier 1 too close to callone pass, not measured

What measurably helps on this job

Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.

Via API

SkillModelBetter than no help byAgainst one plain sentence
skills/code-review-web in the catalogGPT-5 mini+26.7 percentage points (+3.2 to +50.1)could not be told apart from the one-line instruction (−8.9 to +35.5)
Nothing left to measure: these models already scored near the top with no help.
ModelWhereWith no help
Claude Opus 5In Claude Code93.5%, tier 1

How hard the tasks were

These scores are from the first tasks we built for this job, tier 1, the set the models are read on.

The full tables behind this page

Via API

Read from Wider run, three ways, 2026-09-06.

In Claude Code

Read from Wider run, three ways, 2026-09-06, with the same tasks also from Opus 5.5 series, wider run, three ways, 2026-09-26; Fable 5, wider run, three ways, 2026-09-19.

Every result above, with the run it was read from.
SkillModelTest methodRun
skills/code-review-web in the catalogGPT-5 miniVia APIcode-review-tier-skills, tier 3

Version pairs on this job

  • On code review, wider run, Claude Fable 5.1 finds 89% of seeded code defects unaided against 86% for Claude Fable 5; no measurable change.

The job in full: Code review. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/code-review.json.