Does Accessibility help? Tested on accessibility tasks

Accessibility, from affaan-m/ECC. Its accessibility audit tasks were run with the skill loaded and without it. Each time it is the same task, run twice.

What we tested

Whether loading this skill helps a model find accessibility problems on a page.

What counts as helping

The skill has to find at least 10 percentage points more of the planted problems than the model finds without it, on the same 10 test pages.

How we scored it

We planted 48 known problems across the 10 test pages and counted how many the model named. The score is the share it found, from 0 to 1.

Without to with, per model

Via API tested via API

GPT-5 mini First run, with and without the skill

0.705
0.622

No measured effect

GPT-5 mini Wider run, three ways

0.597
0.608

No measured effect

Gemini 3.1 Flash Lite First run, with and without the skill

0.633
0.538

No measured effect

Gemini 3.1 Flash Lite Wider run, three ways

0.673
0.538

No measured effect

Claude Haiku 4.5 First run, with and without the skill

0.390
0.521

No measured effect

deterministic pass rate, 0 to 1

In Claude Code tested in Claude Code

Claude Opus 5 Wider run, three ways

0.877
0.915

No measured effect

Claude Opus 5 Second run, every task twice

0.857
0.880

No measured effect

Claude Opus 5 First run, single attempt

0.890
0.837

No measured effect

Claude Fable 5 Second run, every task twice

0.800
0.900

No measured effect

Claude Fable 5.1 Wider run, three ways

0.935
0.875

No measured effect

Claude Fable 5.1 Fable 5.1, single attempt

0.842
0.850

No measured effect

Claude Sonnet 5 Second run, every task twice

0.811
0.823

No measured effect

Claude Sonnet 5 Wider run, three ways

0.838
0.784

No measured effect

Claude Haiku 4.5 Second run, every task twice

0.452
0.553

No measured effect

Claude Haiku 4.5 Wider run, three ways

0.435
0.513

No measured effect

Claude Fable 5 First run, single attempt

served model mismatch; 2 usable pairs

Withheld

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Grey is the score without. Colour is the score with. The bracket shows how much the difference could move if we ran it again. Rows are ordered by the score with. A model measured at two versions keeps its versions next to each other. The two test methods are reported separately and never averaged. A bracket wider than the axis is drawn to the edge with its cap omitted; the table gives its bounds.
Show per-task detail

Every item, without and with, per model

Via API tested via API

  • Claude Haiku 4.5
  • GPT-5 mini
  • Gemini 3.1 Flash Lite
  • Gemini 3.1 Flash Lite
  • GPT-5 mini

deterministic pass rate, 0 to 1

Hover or focus an item to read its task and every model’s two values.

Task-paired: these item means are what this cell's interval is built from.

The item means behind this plot
Per-item arm means. Without is the unaided arm, with is the treated arm.
ItemClaude Haiku 4.5 GPT-5 mini Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite GPT-5 mini
Task without with without with without with without with without with
a10 0.250 0.250 0.500 0.250 0.250 0.250 0.250 0.250 0.250 0.250
a06 0.400 0.400 0.400 0.400 0.400 0.400 0.400 0.400 0.200 0.400
a04 0.200 0.200 0.000 1.000 0.600 0.600 0.600 0.600 1.000 1.000
a03 0.400 0.600 0.600 0.800 0.200 0.000 0.600 0.000 0.800 0.400
a08 0.000 0.333 0.667 0.333 0.667 0.333 0.667 0.333 0.667 0.333
a09 0.333 0.667 0.667 0.333 0.667 0.667 0.667 0.667 1.000 1.000
a01 0.500 0.375 0.750 0.750 0.750 0.750 0.750 0.750 0.750 0.625
a05 0.400 0.800 0.800 0.800 0.800 0.800 0.800 0.800 0.800 0.800
a07 0.750 0.750 0.750 0.750 1.000 0.750 1.000 0.750 0.750 0.750
a02 0.667 0.833 0.833 0.667 1.000 0.833 1.000 0.833 0.833 0.667
A picture of the per-item means behind the per-model numbers. The tables are the source. Each model draws two lines over the same task set: a dashed line through its unaided scores and a solid line through its treated ones. Items are ordered by the mean unaided score across the models that measured them, lowest first. The two instruments are reported separately and never averaged.

Without to with, per model

In Claude Code tested in Claude Code

Claude Haiku 4.5 Expansion cohort (three arms)

No measured effect

Claude Haiku 4.5 Panel v2 (repeat sampling)

No measured effect

Claude Fable 5 Panel v2 (repeat sampling)

No measured effect

Claude Sonnet 5 Panel v2 (repeat sampling)

No measured effect

Claude Sonnet 5 Expansion cohort (three arms)

No measured effect

Claude Fable 5.1 Fable 5.1, one pass

No measured effect

Claude Opus 5 Panel v2 (repeat sampling)

No measured effect

Claude Opus 5 Expansion cohort (three arms)

No measured effect

Claude Opus 5 Skill cohort (one pass)

No measured effect

Claude Fable 5.1 Expansion cohort (three arms)

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Each row runs from the score without to the score with. The thin bar beneath it is the bracket, and it shows how much the difference could move if we ran it again. The two test methods are reported separately and never averaged. A bracket wider than the axis is drawn to the edge with its end cap omitted; the table gives its bounds.

The Claude Code runs committed their cells as arm means and intervals rather than as per-item values, so this panel is drawn without to with rather than item by item.

What exactly was tested, and how it was scored
Repository
affaan-m/ECC
Path
skills/accessibility
Commit
d8409a4b0813771235555e32e3d8046a73988bfa
Content hash
ab86a50717d67c4f6dc18cc4b506e06784baaddab4f35c351040dbfa11eb35da
Date tested
2026-08-30

The pin is the whole of this skill's identity here. It resolves at https://github.com/affaan-m/ECC/tree/d8409a4b0813771235555e32e3d8046a73988bfa/skills/accessibility, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.

S04-ecc-accessibility

Loading skills/accessibility from affaan-m/ECC improves outputs on accessibility audit tasks.

Pass criterion
With-arm mean coverage exceeds the without-arm by at least 0.10 on the same 10 fixture pages at the same model and settings.
Scale
unit. Same scorer, same ledger, same scale as S03.
Task pairs planned per model
10
Notes
Shares the accessibility-audit task set with S03, so the two rows are measured on identical pages against one answer key.

Injected context tokens

Injected context tokens, per arm
ArmContext charactersInjected context tokens
without00
with6,5841,546

The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.

Verdict per model

Effect per model
ModelReadingOrbitWithout -> withEffect95 percent intervalTask pairsModel version returned
Claude Haiku 4.5Deterministic, scored against a committed answer keyIn free drift unclear. The model scored 39% without it.0.3900 -> 0.5208+0.1308[0.0187, 0.2429]10 pairs, 20 of 20 recordsclaude-haiku-4-5-20251001
Gemini 3.1 Flash LiteDeterministic, scored against a committed answer keyIn free drift unclear. The model scored 63% without it.0.6333 -> 0.5383-0.0950[-0.1753, -0.0147]10 pairs, 20 of 20 recordsgemini-3.1-flash-lite
GPT-5 miniDeterministic, scored against a committed answer keyIn free drift unclear. The model scored 71% without it.0.7050 -> 0.6225-0.0825[-0.1931, 0.0281]10 pairs, 20 of 20 recordsgpt-5-mini-2025-08-07
Gemini 3.1 Flash LiteDeterministic, scored against a committed answer keyIn free drift unclear. The model scored 63% without it.0.6733 -> 0.5383-0.1350[-0.2622, -0.0078]10 pairs, 30 of 30 recordsgemini-3.1-flash-lite
GPT-5 miniDeterministic, scored against a committed answer keyIn free drift (10 of 10 pairs) unclear. The model scored 48% without it.0.5967 -> 0.6083+0.0117[-0.2285, 0.2518]10 pairs, 30 of 31 recordsgpt-5-mini-2025-08-07, gpt-5-mini
  • In free drift: no separation the design can resolve. Not evidence of no effect.

Cost

Cost per task
ModelArmTasks attemptedMean input tokensMean output tokensCost per taskCost ratioGrading cost per task pair
Claude Haiku 4.5 served claude-haiku-4-5-20251001without1060056$0.00088
Claude Haiku 4.5 served claude-haiku-4-5-20251001with102,40297$0.002893.28x
GPT-5 mini served gpt-5-mini-2025-08-07without10515107$0.00034
GPT-5 mini served gpt-5-mini-2025-08-07with102,05499$0.000712.08x
Gemini 3.1 Flash Litewithout1052859$0.00022
Gemini 3.1 Flash Litewith102,16355$0.000622.83x

Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.

Where the with arm did not win

Claude Haiku 4.5, Deterministic, scored against a committed answer key

Lost on 1 of 10 pairs: a01 (-0.1250).

Drew on 4 of 10 pairs: a04, a06, a07, a10.

Findings not in the answer key: 52 with the skill, 24 without. These are counted and never netted off coverage.

Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key

Lost on 4 of 10 pairs: a08 (-0.3333), a07 (-0.2500), a03 (-0.2000), a02 (-0.1667).

Drew on 6 of 10 pairs: a01, a04, a05, a06, a09, a10.

Findings not in the answer key: 19 with the skill, 18 without. These are counted and never netted off coverage.

GPT-5 mini, Deterministic, scored against a committed answer key

Lost on 4 of 10 pairs: a03 (-0.4000), a08 (-0.3333), a02 (-0.1667), a01 (-0.1250).

Drew on 5 of 10 pairs: a04, a05, a07, a09, a10.

Findings not in the answer key: 48 with the skill, 53 without. These are counted and never netted off coverage.

Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key

Lost on 4 of 10 pairs: a03 (-0.6000), a08 (-0.3333), a07 (-0.2500), a02 (-0.1667).

Drew on 6 of 10 pairs: a01, a04, a05, a06, a09, a10.

Findings not in the answer key: 19 with the skill, 18 without. These are counted and never netted off coverage.

GPT-5 mini, Deterministic, scored against a committed answer key

Lost on 4 of 10 pairs: a08 (-0.3333), a09 (-0.3333), a10 (-0.2500), a02 (-0.1667).

Drew on 4 of 10 pairs: a01, a05, a06, a07.

Findings not in the answer key: 47 with the skill, 58 without. These are counted and never netted off coverage.

In Claude Code

These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument: no figure here is averaged with one tested via API above. How the two were compared.

  • Claude Opus 5 · unclear, and both runs agree
  • Claude Fable 5 · unclear under repeat sampling
  • Claude Haiku 4.5 · unclear under repeat sampling
  • Claude Sonnet 5 · unclear under repeat sampling
  • Claude Fable 5.1 · unclear in the version re-test

Skill cohort (one pass)

one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokensFalse positives per page
Claude Opus 50.89000.8367-0.053-0.126 to 0.01910In free drift2,59719 with, 17 without

Verdict withheld on this run

Claude Fable 5: served model mismatch; 2 usable pairs. The plan served a different model on most of one arm, so these records are evidence about that other model and are scored into no cell. They are kept, and their tokens still count against what the run consumed.

Panel v2 (repeat sampling)

two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 50.80000.9000+0.100-0.096 to 0.2962In free drift107,700
Claude Haiku 4.50.45210.5533+0.101-0.045 to 0.24710In free drift66,632
Claude Opus 50.85670.8800+0.023-0.047 to 0.09410In free drift89,772
Claude Sonnet 50.81080.8225+0.012-0.038 to 0.06110In free drift102,812
Thinking tokens, recorded per arm. Reported, never scored: no interval or Orbit on this page consults them.
ModelArmRecords that reasonedMean thinking tokensMost on one record
Claude Fable 5with skill20 of 201,4053,393
Claude Fable 5no skill38 of 38522965

Fable 5.1, one pass

one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokensFalse positives per page
Claude Fable 5.10.84170.8500+0.008-0.078 to 0.09420In free driftnot measured14 with, 12 without

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokensFalse positives per page
Claude Fable 5.10.93500.8750-0.060-0.144 to 0.02410In free driftnot measured: this run records the dose per arm, in the doses list13 with, 11 without
Claude Haiku 4.50.43500.5133+0.078-0.002 to 0.15910In free driftnot measured: this run records the dose per arm, in the doses list33 with, 24 without
Claude Opus 50.87670.9150+0.038-0.054 to 0.13010In free driftnot measured: this run records the dose per arm, in the doses list20 with, 23 without
Claude Sonnet 50.83750.7842-0.053-0.164 to 0.05810In free driftnot measured: this run records the dose per arm, in the doses list57 with, 37 without

Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.

720 run records behind this page. Every verdict is computed at build time by the same frozen status_v1 rule that decides every other verdict on this site, and nothing here is written by hand.