Best LLM for SEO Audits, Tested
An on-page SEO audit means reading a web page and finding the search problems that were put in it. Each model was given the same ten tasks with no help, tested in the API and Claude Code.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
Gemini 3.1 Flash Lite 48.8%
Too close to call with GPT-5 mini and Claude Haiku 4.5
tier 1, First run, with and without the skill In Claude Code
Claude Fable 5.1 97.2%
Too close to call with Claude Opus 5, Claude Opus 5.5 and Claude Fable 5
tier 1, Wider run, three ways
Some of these models come from another run of the same tasks | Via API
Gemini 3.1 Flash Lite $0.00023 a task
The best score is also the cheapest in range
billed through the API In Claude Code
Claude Opus 5 $0.00715 a task
Against Claude Fable 5.1 $0.03873 a task, the best score
run on subscription; shown at API list price for comparison | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 33 of 40
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 23.9% tier 1 too close to call | one pass, not measured |
| GPT-5 mini | 44.9% tier 1 too close to call | one pass, not measured |
| Gemini 3.1 Flash Lite | 48.8% tier 1 best | one pass, not measured |
| GPT-5.4 mini | not run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 23.1% tier 1 | one pass, not measured |
| Claude Fable 5 | 93.2% tier 1 too close to call | one pass, not measured |
| Claude Fable 5.1 | 97.2% tier 1 best | one pass, not measured |
| Claude Opus 5 | 88.1% tier 1 too close to call | one pass, not measured |
| Claude Opus 5.5 | 93.3% tier 1 too close to call | one pass, not measured |
| Claude Sonnet 5 | 53.0% tier 1 | one pass, not measured |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Skill | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| skills/seo-onpage in the catalog | Gemini 3.1 Flash Lite | +21.9 percentage points (+4.6 to +39.2) | the one-line instruction was not asked on this run |
| Model | Where | With no help |
|---|---|---|
| Claude Fable 5 | In Claude Code | 93.2%, tier 1 |
| Claude Fable 5.1 | In Claude Code | 97.2%, tier 1 |
| Claude Opus 5.5 | In Claude Code | 93.3%, tier 1 |
How hard the tasks were
These scores are from the first tasks we built for this job, tier 1, the set the models are read on.
The full tables behind this page
Via API
Read from First run, with and without the skill, 2026-08-30.
- Another run, Wider run, three ways, 2026-09-06: GPT-5 mini 52.2% (tier 1), Gemini 3.1 Flash Lite 48.8% (tier 1)
In Claude Code
Read from Wider run, three ways, 2026-09-06, with the same tasks also from Opus 5.5 series, wider run, three ways, 2026-09-26; Fable 5, wider run, three ways, 2026-09-19.
- Another run, Second run, every task twice, 2026-08-31: Claude Opus 5 88.5% (tier 1), Claude Fable 5 88.0% (tier 1), Claude Sonnet 5 59.0% (tier 1), Claude Haiku 4.5 23.1% (tier 1)
- Another run, First run, single attempt, 2026-08-31: Claude Fable 5 92.3% (tier 1), Claude Opus 5 87.9% (tier 1)
- Another run, Fable 5.1, single attempt, 2026-09-03: Claude Fable 5.1 98.3% (tier 1)
| Skill | Model | Test method | Run |
|---|---|---|---|
| skills/seo-onpage in the catalog | Gemini 3.1 Flash Lite | Via API | First run, with and without the skill, tier 1 |
Version pairs on this job
- On on-page audit, Claude Fable 5.1 finds 98% of seeded on-page issues unaided against 92% for Claude Fable 5; scores higher.
- On on-page audit, wider run, Claude Fable 5.1 finds 97% of seeded on-page issues unaided against 93% for Claude Fable 5; no measurable change.
The job in full: On-page audit. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/on-page-audit.json.