Best LLM for Code Review, Tested
A code review means reading a change to some code and finding the mistakes that were put in it. Each model was given the same ten tasks with no help, tested in the API and Claude Code.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
GPT-5 mini 70.0%
Too close to call with Gemini 3.1 Flash Lite
tier 1, Wider run, three ways In Claude Code
Claude Opus 5 93.5%
Too close to call with Claude Fable 5.1, Claude Sonnet 5, Claude Opus 5.5 and Claude Fable 5
tier 1, Wider run, three ways
Some of these models come from another run of the same tasks | Via API
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. In Claude Code
Claude Sonnet 5 $0.00406 a task
Against Claude Opus 5 $0.00948 a task, the best score
run on subscription; shown at API list price for comparison | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | not run | |
| GPT-5 mini | 70.0% tier 1 best | one pass, not measured |
| Gemini 3.1 Flash Lite | 57.5% tier 1 too close to call | one pass, not measured |
| GPT-5.4 mini | not run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 58.3% tier 1 | one pass, not measured |
| Claude Fable 5 | 85.8% tier 1 too close to call | one pass, not measured |
| Claude Fable 5.1 | 89.0% tier 1 too close to call | one pass, not measured |
| Claude Opus 5 | 93.5% tier 1 best | one pass, not measured |
| Claude Opus 5.5 | 85.4% tier 1 too close to call | one pass, not measured |
| Claude Sonnet 5 | 74.9% tier 1 too close to call | one pass, not measured |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Skill | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| skills/code-review-web in the catalog | GPT-5 mini | +26.7 percentage points (+3.2 to +50.1) | could not be told apart from the one-line instruction (−8.9 to +35.5) |
| Model | Where | With no help |
|---|---|---|
| Claude Opus 5 | In Claude Code | 93.5%, tier 1 |
How hard the tasks were
These scores are from the first tasks we built for this job, tier 1, the set the models are read on.
The full tables behind this page
Via API
Read from Wider run, three ways, 2026-09-06.
In Claude Code
Read from Wider run, three ways, 2026-09-06, with the same tasks also from Opus 5.5 series, wider run, three ways, 2026-09-26; Fable 5, wider run, three ways, 2026-09-19.
| Skill | Model | Test method | Run |
|---|---|---|---|
| skills/code-review-web in the catalog | GPT-5 mini | Via API | code-review-tier-skills, tier 3 |
Version pairs on this job
- On code review, wider run, Claude Fable 5.1 finds 89% of seeded code defects unaided against 86% for Claude Fable 5; no measurable change.
The job in full: Code review. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/code-review.json.