engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills
Accessibility audit tasks, run with the skill loaded and without it, on the same task pairs.
Identity pin
- Repository
- alirezarezvani/claude-skills
- Path
- engineering-team/a11y-audit/skills/a11y-audit
- Commit
19392f7a08264ed00486a251f5b2098321771f94- Content hash
34f619aeb9ccbfc2f09cd0b0188ae2e4d99f87e9d78c77a16acb2f7e6c450d01- Date tested
- not yet tested
The pin is the whole of this skill's identity here. It resolves at https://github.com/alirezarezvani/claude-skills/tree/19392f7a08264ed00486a251f5b2098321771f94/engineering-team/a11y-audit/skills/a11y-audit, and the content hash is a sha256 over exactly the text the with arm was given. Nothing else about the skill appears on this site.
What was claimed
- Loading engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills improves outputs on accessibility audit tasks.
Pass criterion: With-arm mean coverage exceeds the without-arm by at least 0.10 on the same 10 fixture pages at the same model and settings. Scale: unit. Same scorer, same instrument and same scale as S04-ecc-accessibility. Added 2026-08-31 by the source expansion under R23; adding a row to an instrument does not touch the instrument.
Shares the accessibility-audit task set with every other claim in this class, so all of its rows are measured on identical items against one answer key. The slot this claim scores was pinned by the 2026-08-31 source expansion under R23 and NO MODEL HAS BEEN RUN AGAINST IT. The claim exists so the class holds one claim per pinned slot, which is what makes the row comparable the day it is run.
Injected context tokens
| Arm | Context characters | Injected context tokens |
|---|---|---|
| without | 0 | 0 |
| with | 56,350 | 13,228 |
The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.
Not yet measured
No run records exist for this skill. The instrument is built and committed, the task sets are frozen, and the projected cost of the run is published, but no model has been called. There is therefore no verdict, no effect, no interval and no cost per task on this page, and none is shown.
- S15-alirezarezvani-a11y-audit: 10 task pairs per model, as planned. Deterministic, scored against a committed answer key.
The committed instrument is at harness/claims/claims-skills.json, harness/claims/tasksets-skills.json and harness/claims/skill-cohort.json. The projection for the run that has not happened is at harness/results/dry-run-projection-skills.md. How a claim gets tested.