Which LLM Finds Accessibility Problems? Tested

An accessibility audit means reading a web page and finding the problems that would stop someone using it. Each model was given the same 30 tasks with no help, tested in the API and Claude Code.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API Each model was given tasks at its own level of difficulty, so no one model is ranked first.
In Claude Code Each model was given tasks at its own level of difficulty, so no one model is ranked first.
Via API Each model was given tasks at its own level of difficulty, so no one model is ranked first.
In Claude Code Each model was given tasks at its own level of difficulty, so no one model is ranked first.
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code Claude Opus 5, same answer on 38 of 40 no sampling control here, Second run, every task twice
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.5in another run
GPT-5 mini58.7% tier 1not asked twice
Gemini 3.1 Flash Lite40.0% tier 2not asked twice
GPT-5.4 mininot run

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.544.8% tier 1not asked twice
Claude Fable 5in another run
Claude Fable 5.189.4% tier 3 too close to callnot asked twice
Claude Opus 556.1% tier 3not asked twice
Claude Opus 5.592.8% tier 3 bestnot asked twice
Claude Sonnet 543.3% tier 2not asked twice

What measurably helps on this job

Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.

Via API

SkillModelBetter than no help byAgainst one plain sentence
skills/accessibility-audit in the catalogClaude Haiku 4.5+28.9 percentage points (+14.5 to +43.3)the one-line instruction was not asked on this run
plugins/accessibility-compliance/skills/wcag-audit-patternsGPT-5 mini+23.5 percentage points (+5.4 to +41.6)could not be told apart from the one-line instruction (−2.7 to +29.2)

In Claude Code

SkillModelBetter than no help byAgainst one plain sentence
skills/accessibility-audit in the catalogClaude Haiku 4.5+23.2 percentage points (+8.8 to +37.7)the one-line instruction was not asked on this run
engineering-team/a11y-audit/skills/a11y-auditClaude Haiku 4.562.4% with it against 48.0% with no helpbeat the one-line instruction by +20.5 percentage points (+9.7 to +31.3)
Nothing left to measure: these models already scored near the top with no help.
ModelWhereWith no help
Claude Opus 5.5In Claude Code92.8%, tier 3

How hard the tasks were

Each model takes its next skills test at its own tier: the one where its score, with no skill loaded, leaves room to improve. A model that scores high everywhere gets the hardest tier. So a skill's result can be compared with other skills on the same model, but not used to rank one model against another, because the models did different work.

The full tables behind this page

Via API

Read from Harder tasks, each model at its own tier, 2026-09-18.

  • Another run, First run, with and without the skill, 2026-08-30: GPT-5 mini 69.9% (tier 1), Gemini 3.1 Flash Lite 63.3% (tier 1), Claude Haiku 4.5 40.0% (tier 1)
  • Another run, Wider run, three ways, 2026-09-06: Gemini 3.1 Flash Lite 63.3% (tier 1), GPT-5 mini 58.2% (tier 1)

In Claude Code

Read from Harder tasks, each model at its own tier, 2026-09-18, with the same tasks also from Harder tasks, Claude Opus 5.5 at its own tier, 2026-09-26.

  • Another run, Wider run, three ways, 2026-09-06: Claude Fable 5.1 87.9% (tier 1), Claude Opus 5 87.1% (tier 1), Claude Sonnet 5 80.2% (tier 1), Claude Haiku 4.5 46.1% (tier 1)
  • Another run, Second run, every task twice, 2026-08-31: Claude Opus 5 87.3% (tier 1), Claude Sonnet 5 82.3% (tier 1), Claude Haiku 4.5 45.5% (tier 1)
  • Another run, Opus 5.5 series, wider run, three ways, 2026-09-26: Claude Opus 5.5 85.8% (tier 1)
  • Another run, Fable 5, wider run, three ways, 2026-09-19: Claude Fable 5 89.7% (tier 1)
  • Another run, Fable 5.1, single attempt, 2026-09-03: Claude Fable 5.1 85.8% (tier 1)
  • Another run, First run, single attempt, 2026-08-31: Claude Opus 5 89.0% (tier 1)
Every result above, with the run it was read from.
SkillModelTest methodRun
skills/accessibility-audit in the catalogClaude Haiku 4.5Via APIFirst run, with and without the skill, tier 1
skills/accessibility-audit in the catalogClaude Haiku 4.5In Claude CodeSecond run, every task twice, tier 1
engineering-team/a11y-audit/skills/a11y-auditClaude Haiku 4.5In Claude CodeWider run, three ways, tier 1
plugins/accessibility-compliance/skills/wcag-audit-patternsGPT-5 miniVia APIaccessibility-tier-skills, tier 1

Version pairs on this job

  • On accessibility audit, wider run, Claude Fable 5.1 finds 88% of seeded accessibility defects unaided against 90% for Claude Fable 5; no measurable change.

The job in full: Accessibility audit. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/accessibility-audit.json.