Which LLM Finds Accessibility Problems? Tested
An accessibility audit means reading a web page and finding the problems that would stop someone using it. Each model was given the same 30 tasks with no help, tested in the API and Claude Code.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
Each model was given tasks at its own level of difficulty, so no one model is ranked first. In Claude Code
Each model was given tasks at its own level of difficulty, so no one model is ranked first. | Via API
Each model was given tasks at its own level of difficulty, so no one model is ranked first. In Claude Code
Each model was given tasks at its own level of difficulty, so no one model is ranked first. | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 38 of 40
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | in another run | |
| GPT-5 mini | 58.7% tier 1 | not asked twice |
| Gemini 3.1 Flash Lite | 40.0% tier 2 | not asked twice |
| GPT-5.4 mini | not run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 44.8% tier 1 | not asked twice |
| Claude Fable 5 | in another run | |
| Claude Fable 5.1 | 89.4% tier 3 too close to call | not asked twice |
| Claude Opus 5 | 56.1% tier 3 | not asked twice |
| Claude Opus 5.5 | 92.8% tier 3 best | not asked twice |
| Claude Sonnet 5 | 43.3% tier 2 | not asked twice |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Skill | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| skills/accessibility-audit in the catalog | Claude Haiku 4.5 | +28.9 percentage points (+14.5 to +43.3) | the one-line instruction was not asked on this run |
| plugins/accessibility-compliance/skills/wcag-audit-patterns | GPT-5 mini | +23.5 percentage points (+5.4 to +41.6) | could not be told apart from the one-line instruction (−2.7 to +29.2) |
In Claude Code
| Skill | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| skills/accessibility-audit in the catalog | Claude Haiku 4.5 | +23.2 percentage points (+8.8 to +37.7) | the one-line instruction was not asked on this run |
| engineering-team/a11y-audit/skills/a11y-audit | Claude Haiku 4.5 | 62.4% with it against 48.0% with no help | beat the one-line instruction by +20.5 percentage points (+9.7 to +31.3) |
| Model | Where | With no help |
|---|---|---|
| Claude Opus 5.5 | In Claude Code | 92.8%, tier 3 |
How hard the tasks were
Each model takes its next skills test at its own tier: the one where its score, with no skill loaded, leaves room to improve. A model that scores high everywhere gets the hardest tier. So a skill's result can be compared with other skills on the same model, but not used to rank one model against another, because the models did different work.
The full tables behind this page
Via API
Read from Harder tasks, each model at its own tier, 2026-09-18.
- Another run, First run, with and without the skill, 2026-08-30: GPT-5 mini 69.9% (tier 1), Gemini 3.1 Flash Lite 63.3% (tier 1), Claude Haiku 4.5 40.0% (tier 1)
- Another run, Wider run, three ways, 2026-09-06: Gemini 3.1 Flash Lite 63.3% (tier 1), GPT-5 mini 58.2% (tier 1)
In Claude Code
Read from Harder tasks, each model at its own tier, 2026-09-18, with the same tasks also from Harder tasks, Claude Opus 5.5 at its own tier, 2026-09-26.
- Another run, Wider run, three ways, 2026-09-06: Claude Fable 5.1 87.9% (tier 1), Claude Opus 5 87.1% (tier 1), Claude Sonnet 5 80.2% (tier 1), Claude Haiku 4.5 46.1% (tier 1)
- Another run, Second run, every task twice, 2026-08-31: Claude Opus 5 87.3% (tier 1), Claude Sonnet 5 82.3% (tier 1), Claude Haiku 4.5 45.5% (tier 1)
- Another run, Opus 5.5 series, wider run, three ways, 2026-09-26: Claude Opus 5.5 85.8% (tier 1)
- Another run, Fable 5, wider run, three ways, 2026-09-19: Claude Fable 5 89.7% (tier 1)
- Another run, Fable 5.1, single attempt, 2026-09-03: Claude Fable 5.1 85.8% (tier 1)
- Another run, First run, single attempt, 2026-08-31: Claude Opus 5 89.0% (tier 1)
| Skill | Model | Test method | Run |
|---|---|---|---|
| skills/accessibility-audit in the catalog | Claude Haiku 4.5 | Via API | First run, with and without the skill, tier 1 |
| skills/accessibility-audit in the catalog | Claude Haiku 4.5 | In Claude Code | Second run, every task twice, tier 1 |
| engineering-team/a11y-audit/skills/a11y-audit | Claude Haiku 4.5 | In Claude Code | Wider run, three ways, tier 1 |
| plugins/accessibility-compliance/skills/wcag-audit-patterns | GPT-5 mini | Via API | accessibility-tier-skills, tier 1 |
Version pairs on this job
- On accessibility audit, wider run, Claude Fable 5.1 finds 88% of seeded accessibility defects unaided against 90% for Claude Fable 5; no measurable change.
The job in full: Accessibility audit. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/accessibility-audit.json.