Claude skills for Accessibility, tested

Skills for accessibility work, measured on accessibility audit tasks. Each was the same task, run twice, once with the skill loaded and once without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.

Measured on 3 models (Claude, GPT & Gemini): Claude Haiku 4.5, GPT-5 mini, Gemini 3.1 Flash Lite.

10 fixture web pages, each seeded with known WCAG defects, 48 in total. The model is asked to list the defects it finds; the score is the share of seeded defects it names. 23 defects are verified in the fixtures by mechanical check; 25 are judgement items the check cannot verify and are listed as such.

What helped, and on how many models

  • First

    skills/accessibility-audit rampstackco/claude-skills*

    Accessibility audit

    Helped on Claude Haiku 4.5; no reliable difference on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.

    Found 71% of the seeded accessibility defects, against 58% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · helped
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.0063. Adds about 11,700 tokens of context to every request.

  • Tied for second · shared with 4 others, and the tie is broken by nothing

    skills/accessibility-audit rampstackco/claude-skills*

    Accessibility audit

    No reliable difference on any model tested. Measured against the same task with a one-line instruction.

    Found 68% of the seeded accessibility defects, against 60% with no skill loaded. Against 60% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · helped
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.003. Adds about 11,300 tokens of context to every request.

  • Tied for second · shared with 4 others, and the tie is broken by nothing

    plugins/accessibility-compliance/skills/wcag-audit-patterns wshobson/agents

    Accessibility audit

    No reliable difference on any model tested. Measured against the same task with a one-line instruction.

    Found 63% of the seeded accessibility defects, against 63% with no skill loaded. Against 58% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00113. Adds about 4,000 tokens of context to every request.

  • Tied for second · shared with 4 others, and the tie is broken by nothing

    engineering-team/a11y-audit/skills/a11y-audit alirezarezvani/claude-skills

    Accessibility audit

    No reliable difference on any model tested. Measured against the same task with a one-line instruction.

    Found 65% of the seeded accessibility defects, against 64% with no skill loaded. Against 62% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · helped
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00424. Adds about 16,400 tokens of context to every request.

  • Tied for second · shared with 4 others, and the tie is broken by nothing

    skills/accessibility affaan-m/ECC

    Accessibility audit

    No reliable difference on any model tested. Measured against no skill loaded.

    Found 56% of the seeded accessibility defects, against 58% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · unclear
    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · unclear
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00141. Adds about 2,200 tokens of context to every request.

  • Tied for second · shared with 4 others, and the tie is broken by nothing

    skills/accessibility affaan-m/ECC

    Accessibility audit

    No reliable difference on any model tested. Measured against the same task with a one-line instruction.

    Found 57% of the seeded accessibility defects, against 55% with no skill loaded. Against 64% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · unclear
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.000669. Adds about 2,100 tokens of context to every request.

The numbers

How to read these tables

How to read these tables

Without -> with
The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
Effect (range)
With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
Orbit
  • Stable: delta at or above the pass threshold, and the interval excludes zero
  • Past the horizon: delta at or below the failure floor, and the interval excludes zero
  • In free drift: the interval spans zero, or the delta sits between the floor and the threshold
  • Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound

A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.

Cohort
The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
Commit
The exact version tested. Later commits are not covered by this verdict.
Injected context tokens
How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
Tested
The date. Verdicts age as models change.

Accessibility audit

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

engineering-team/a11y-audit/skills/a11y-audit Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.763. Instruction only: 0.673. Change +0.130. It could sit between -0.018 and +0.278.

skills/accessibility-audit Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.728. Instruction only: 0.673. Change +0.094. It could sit between -0.068 and +0.257.

skills/accessibility-audit Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.728. Change +0.094. It could sit between -0.068 and +0.257.

skills/accessibility-audit Claude Haiku 4.5 First run, with and without the skill, Deterministic

Holds

No skill: 0.410. With the skill: 0.699. Change +0.289. It could sit between +0.145 and +0.433.

plugins/accessibility-compliance/skills/wcag-audit-patterns Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.690. Instruction only: 0.673. Change +0.057. It could sit between -0.085 and +0.199.

skills/accessibility-audit GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.689. Change -0.003. It could sit between -0.112 and +0.105.

skills/accessibility-audit GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.626. Instruction only: 0.529. Change +0.051. It could sit between -0.074 and +0.175.

skills/accessibility GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.705. With the skill: 0.622. Change -0.083. It could sit between -0.193 and +0.028.

skills/accessibility GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.476. With the skill: 0.608. Instruction only: 0.597. Change +0.133. It could sit between -0.030 and +0.295.

plugins/accessibility-compliance/skills/wcag-audit-patterns GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.635. With the skill: 0.575. Instruction only: 0.488. Change -0.060. It could sit between -0.198 and +0.078.

engineering-team/a11y-audit/skills/a11y-audit GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.641. With the skill: 0.542. Instruction only: 0.568. Change -0.099. It could sit between -0.230 and +0.032.

skills/accessibility Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.538. Instruction only: 0.673. Change -0.095. It could sit between -0.175 and -0.015.

skills/accessibility Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.538. Change -0.095. It could sit between -0.175 and -0.015.

skills/accessibility Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 0.390. With the skill: 0.521. Change +0.131. It could sit between +0.019 and +0.243.

In Claude Code tested in Claude Code

skills/accessibility-audit Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.875. With the skill: 0.935. Change +0.060. It could sit between -0.024 and +0.144.

engineering-team/a11y-audit/skills/a11y-audit Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.895. With the skill: 0.915. Change +0.020. It could sit between -0.071 and +0.111.

skills/accessibility Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.877. With the skill: 0.915. Change +0.038. It could sit between -0.054 and +0.130.

skills/accessibility Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.800. With the skill: 0.900. Change +0.100. It could sit between -0.096 and +0.296.

skills/accessibility-audit Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.855. With the skill: 0.895. Change +0.040. It could sit between -0.012 and +0.092.

engineering-team/a11y-audit/skills/a11y-audit Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.882. With the skill: 0.890. Change +0.008. It could sit between -0.078 and +0.094.

skills/accessibility Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.857. With the skill: 0.880. Change +0.023. It could sit between -0.047 and +0.094.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.895. With the skill: 0.875. Change -0.020. It could sit between -0.059 and +0.019.

skills/accessibility Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.875. Change -0.060. It could sit between -0.144 and +0.024.

skills/accessibility-audit Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.890. With the skill: 0.870. Change -0.020. It could sit between -0.143 and +0.103.

skills/accessibility Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.842. With the skill: 0.850. Change +0.008. It could sit between -0.078 and +0.094.

skills/accessibility-audit Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.890. With the skill: 0.840. Change -0.050. It could sit between -0.178 and +0.078.

skills/accessibility Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.890. With the skill: 0.837. Change -0.053. It could sit between -0.126 and +0.019.

skills/accessibility-audit Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.910. With the skill: 0.830. Change -0.080. It could sit between -0.200 and +0.040.

skills/accessibility Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.811. With the skill: 0.823. Change +0.012. It could sit between -0.038 and +0.061.

engineering-team/a11y-audit/skills/a11y-audit Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.809. With the skill: 0.807. Change -0.002. It could sit between -0.087 and +0.084.

skills/accessibility-audit Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.838. With the skill: 0.804. Change -0.033. It could sit between -0.111 and +0.044.

skills/accessibility-audit Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.800. With the skill: 0.800. Change +0.000. It could sit between +0.000 and +0.000.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.857. With the skill: 0.790. Change -0.067. It could sit between -0.214 and +0.080.

skills/accessibility Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.838. With the skill: 0.784. Change -0.053. It could sit between -0.164 and +0.058.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.818. With the skill: 0.779. Change -0.038. It could sit between -0.163 and +0.087.

skills/accessibility-audit Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.835. With the skill: 0.777. Change -0.058. It could sit between -0.140 and +0.025.

skills/accessibility-audit Claude Haiku 4.5 Second run, every task twice, Deterministic

Holds

No skill: 0.458. With the skill: 0.690. Change +0.232. It could sit between +0.088 and +0.377.

skills/accessibility-audit Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.455. With the skill: 0.628. Change +0.173. It could sit between +0.030 and +0.317.

engineering-team/a11y-audit/skills/a11y-audit Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.419. With the skill: 0.624. Change +0.205. It could sit between +0.097 and +0.313.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.477. With the skill: 0.608. Change +0.131. It could sit between +0.021 and +0.240.

skills/accessibility Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.452. With the skill: 0.553. Change +0.101. It could sit between -0.045 and +0.247.

skills/accessibility Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.435. With the skill: 0.513. Change +0.078. It could sit between -0.002 and +0.159.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.
SkillCohortWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loaded First run, with and without the skillbaseline58%0 by definition02026-08-30
One-line instruction instead First run, with and without the skillbaselineinstruction arm: not measured0 by definition02026-08-30
skills/accessibility-audit from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this testrampstackco/claude-skills*58% -> 71%+0.13 [-0.00, +0.29]In free drift on 2 of 311,7292026-08-30
skills/accessibility from affaan-m/ECC d8409a4b0813 First run, with and without the skillaffaan-m/ECC58% -> 56%-0.02 [-0.10, +0.13]In free drift2,2062026-08-30
No skill loaded Wider run, three waysbaseline61%0 by definition02026-09-05
One-line instruction instead Wider run, three waysbaseline61%0 by definition02026-09-05
skills/accessibility-audit from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this testrampstackco/claude-skills*60% -> 68%+0.08 [+0.05, +0.10]In free drift11,2952026-09-05
plugins/accessibility-compliance/skills/wcag-audit-patterns from wshobson/agents 38e19c20d2b1 Wider run, three wayswshobson/agents63% -> 63%+0.05 [+0.02, +0.09]In free drift3,9672026-09-05
engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills 19392f7a0826 Wider run, three waysalirezarezvani/claude-skills64% -> 65%+0.03 [-0.03, +0.09]In free drift16,4072026-09-05
skills/accessibility from affaan-m/ECC d8409a4b0813 Wider run, three waysaffaan-m/ECC55% -> 57%-0.06 [-0.13, +0.01]In free drift2,1082026-09-05
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.