Claude Skills Compared, Task by Task

Each skill here was tested on a fixed task set, once with the skill loaded into the model's context and once without. The tables show what changed. Nothing on this page describes what a skill is or says about itself; the commit link goes to the source.

How to read this

  • Each skill was run on the same tasks with the skill loaded and without it, and where the third arm was run, with a one-line instruction naming the task instead.
  • The sentence says what happened, against which baseline, and on which models.
  • The bar shows what the model managed without the skill, and what the skill added.
  • Loading a skill costs context on every request, whether or not it helps.
How to read the numbers

How to read these tables

Without the skill
The score on the task set before any skill was loaded. Where this is already near the top of the scale, there was little room for a skill to show an effect, and the result says so instead of reporting no difference.
Without -> with
The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
Effect (range)
With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
Orbit
  • Stable: delta at or above the pass threshold, and the interval excludes zero
  • Past the horizon: delta at or below the failure floor, and the interval excludes zero
  • In free drift: the interval spans zero, or the delta sits between the floor and the threshold
  • Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound

A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.

Cohort
The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
Commit
The exact version tested. Later commits are not covered by this verdict.
Injected context tokens
How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
Tested
The date. Verdicts age as models change.

Skill authoring

The model is given a short brief and asked to write a complete SKILL.md. The score is the share of the open Agent Skills specification's requirements the result meets, counting only the 14 defined requirements that apply to the document written; requirements for optional fields the document does not use are not counted. A second reading against this site's own house style is shown on each skill's page but is not ranked, because one of the skills tested is ours.

5 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.

Every skill tested for writing skills, ranked.

Skill cohort (launch cells)

one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.

  • Helped on someBest result in this test

    skills/skill-creation-walkthrough rampstackco/claude-skills*

    Helped clearly on Claude Haiku 4.5 and GPT-5 mini; helped a little on Gemini 3.1 Flash Lite, under our bar. Measured against no skill loaded.

    Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · unclear. The model scored 87% without it.
    • GPT-5 mini · helped

    Adds about 8,400 tokens of context to every request.

  • Helped on some

    skills/skill-creator anthropics/skills

    Helped clearly on Claude Haiku 4.5 and GPT-5 mini; helped a little on Gemini 3.1 Flash Lite, under our bar. Measured against no skill loaded. Based on 56 of 60 runs; the rest returned nothing.

    Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · unclear. The model scored 87% without it.
    • GPT-5 mini · helped

    Adds about 11,500 tokens of context to every request.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.

  • Helped clearlyBest result in this test

    skills/skill-creation-walkthrough rampstackco/claude-skills*

    Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction. Based on 60 of 63 runs; the rest returned nothing.

    Found 100% of the specification requirements, against 76% with no skill loaded. Against 68% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    Adds about 8,200 tokens of context to every request.

  • Helped clearly

    skills/skill-creator anthropics/skills

    Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction. Based on 60 of 112 runs; the rest returned nothing.

    Found 100% of the specification requirements, against 82% with no skill loaded. Against 70% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    Adds about 10,900 tokens of context to every request.

  • Helped clearly

    distribution/claude-plugin/skills/skill-builder yusufkaraaslan/Skill_Seekers

    Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction. Based on 60 of 67 runs; the rest returned nothing.

    Found 99% of the specification requirements, against 70% with no skill loaded. Against 73% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    Adds about 1,200 tokens of context to every request.

  • Helped on some

    skills/skill-creator openclaw/openclaw

    Helped on Gemini 3.1 Flash Lite; helped a little on GPT-5 mini, under our bar. Measured against the same task with a one-line instruction. Based on 60 of 74 runs; the rest returned nothing.

    Found 99% of the specification requirements, against 75% with no skill loaded. Against 77% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · unclear. The model scored 63% without it.

    Adds about 600 tokens of context to every request.

Show the numbers
SkillCohortWithout the skillWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loaded Wider run, three waysbaseline76%76%0 by definition02026-09-05
One-line instruction instead Wider run, three waysbaseline72%0 by definition02026-09-05
skills/skill-creation-walkthrough from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this testrampstackco/claude-skills*76%76% -> 100%+0.32 [+0.25, +0.38]Stable8,1902026-09-05
skills/skill-creator from anthropics/skills 3b3fad96af16 Wider run, three waysanthropics/skills82%82% -> 100%+0.30 [+0.25, +0.34]Stable10,8722026-09-05
distribution/claude-plugin/skills/skill-builder from yusufkaraaslan/Skill_Seekers f3972efa33fa Wider run, three waysyusufkaraaslan/Skill_Seekers70%70% -> 99%+0.26 [+0.25, +0.27]Stable1,2032026-09-05
skills/skill-creator from openclaw/openclaw 61b08f0ebb7a Wider run, three waysopenclaw/openclaw75%75% -> 99%+0.22 [+0.19, +0.25]Mixed, see page5872026-09-05
No skill loaded First run, with and without the skillbaseline74%74%0 by definition02026-08-30
One-line instruction instead First run, with and without the skillbaselineinstruction arm: not measured0 by definition02026-08-30
skills/skill-creation-walkthrough from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this testrampstackco/claude-skills*74%74% -> 100%+0.26 [+0.13, +0.38]Stable on 2 of 38,3902026-08-30
skills/skill-creator from anthropics/skills 3b3fad96af16 First run, with and without the skillanthropics/skills74%74% -> 100%+0.26 [+0.13, +0.38]Stable on 2 of 311,5142026-08-30
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

distribution/claude-plugin/skills/skill-builder Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creation-walkthrough Claude Haiku 4.5 First run, with and without the skill, Deterministic

Holds

No skill: 0.618. With the skill: 1.000. Change +0.382. It could sit between +0.287 and +0.477.

skills/skill-creation-walkthrough Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creation-walkthrough Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.873. With the skill: 1.000. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creation-walkthrough GPT-5 mini Wider run, three ways, Deterministic

Holds

No skill: 0.655. With the skill: 1.000. Instruction only: 0.618. Change +0.345. It could sit between +0.176 and +0.515.

skills/skill-creation-walkthrough GPT-5 mini First run, with and without the skill, Deterministic

Holds

No skill: 0.718. With the skill: 1.000. Change +0.282. It could sit between +0.128 and +0.435.

skills/skill-creator Claude Haiku 4.5 First run, with and without the skill, Deterministic

Holds

No skill: 0.618. With the skill: 1.000. Change +0.382. It could sit between +0.287 and +0.477.

skills/skill-creator Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creator Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creator Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.873. With the skill: 1.000. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creator GPT-5 mini Wider run, three ways, Deterministic

Holds

No skill: 0.773. With the skill: 1.000. Instruction only: 0.664. Change +0.227. It could sit between +0.105 and +0.350.

skills/skill-creator GPT-5 mini First run, with and without the skill, Deterministic

Holds

No skill: 0.727. With the skill: 1.000. Change +0.273. It could sit between +0.083 and +0.462.

skills/skill-creator GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.627. With the skill: 0.983. Instruction only: 0.791. Change +0.356. It could sit between +0.198 and +0.514.

distribution/claude-plugin/skills/skill-builder GPT-5 mini Wider run, three ways, Deterministic

Holds

No skill: 0.536. With the skill: 0.982. Instruction only: 0.709. Change +0.445. It could sit between +0.268 and +0.623.

In Claude Code tested in Claude Code

distribution/claude-plugin/skills/skill-builder Claude Fable 5.1 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

distribution/claude-plugin/skills/skill-builder Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.600. With the skill: 1.000. Change +0.400. It could sit between +0.293 and +0.507.

distribution/claude-plugin/skills/skill-builder Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

distribution/claude-plugin/skills/skill-builder Claude Sonnet 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.545. With the skill: 1.000. Change +0.455. It could sit between +0.455 and +0.455.

skills/skill-creation-walkthrough Claude Haiku 4.5 Second run, every task twice, Deterministic

Holds

No skill: 0.691. With the skill: 1.000. Change +0.309. It could sit between +0.220 and +0.398.

skills/skill-creation-walkthrough Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Opus 5 Second run, every task twice, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Opus 5 First run, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Fable 5 Second run, every task twice, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Fable 5 First run, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Fable 5.1 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.564. With the skill: 1.000. Change +0.436. It could sit between +0.353 and +0.520.

skills/skill-creator Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.655. With the skill: 1.000. Change +0.345. It could sit between +0.237 and +0.454.

skills/skill-creator Claude Haiku 4.5 Second run, every task twice, Deterministic

Holds

No skill: 0.709. With the skill: 1.000. Change +0.291. It could sit between +0.227 and +0.355.

skills/skill-creator Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Opus 5 First run, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Sonnet 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Sonnet 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.927. Change -0.073. It could sit between -0.168 and +0.022.

skills/skill-creator Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.927. Change -0.073. It could sit between -0.168 and +0.022.

skills/skill-creation-walkthrough Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.927. Change -0.073. It could sit between -0.215 and +0.070.

skills/skill-creator Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.909. Change -0.091. It could sit between -0.269 and +0.087.

skills/skill-creation-walkthrough Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.891. Change -0.109. It could sit between -0.218 and -0.000.

skills/skill-creation-walkthrough Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.855. Change -0.145. It could sit between -0.336 and +0.045.

skills/skill-creation-walkthrough Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.855. Change -0.145. It could sit between -0.336 and +0.045.

skills/skill-creator Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

Scored worse

No skill: 1.000. With the skill: 0.782. Change -0.218. It could sit between -0.436 and -0.000.

skills/skill-creator Claude Fable 5.1 Wider run, three ways, Deterministic

Scored worse

No skill: 1.000. With the skill: 0.564. Change -0.436. It could sit between -0.669 and -0.204.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.

In Claude Code

These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.

Skill cohort (one pass)

one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.

  • skills/skill-creation-walkthrough

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · could not be measured, and both runs agree
  • skills/skill-creator

    • Claude Fable 5 · could not be measured, and both runs agree
    • Claude Opus 5 · unclear under repeat sampling; could not be measured in the single-pass run

Panel v2 (repeat sampling)

two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.

  • skills/skill-creation-walkthrough

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · could not be measured, and both runs agree
    • Claude Haiku 4.5 · helped under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
  • skills/skill-creator

    • Claude Fable 5 · could not be measured, and both runs agree
    • Claude Opus 5 · unclear under repeat sampling; could not be measured in the single-pass run
    • Claude Haiku 4.5 · helped under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling

Fable 5.1, one pass

one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

  • skills/skill-creation-walkthrough

    • Claude Opus 5 · could not be measured, and both runs agree
    • Claude Haiku 4.5 · helped under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · could not be measured in the version re-test
  • skills/skill-creator

    • Claude Opus 5 · unclear under repeat sampling; could not be measured in the single-pass run
    • Claude Haiku 4.5 · helped under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · made results worse in the version re-test
  • skills/skill-creator

    • Claude Fable 5.1 · could not be measured in the three-arm cohort
    • Claude Haiku 4.5 · helped in the three-arm cohort
    • Claude Opus 5 · could not be measured in the three-arm cohort
    • Claude Sonnet 5 · could not be measured in the three-arm cohort
  • distribution/claude-plugin/skills/skill-builder

    • Claude Fable 5.1 · could not be measured in the three-arm cohort
    • Claude Haiku 4.5 · helped in the three-arm cohort
    • Claude Opus 5 · could not be measured in the three-arm cohort
    • Claude Sonnet 5 · could not be measured in the three-arm cohort
Skill authoring · tested via API
SkillClaude Haiku 4.5 · skill-cohort · Gemini 3.1 Flash Lite · skill-cohort · GPT-5 mini · skill-cohort · Gemini 3.1 Flash Lite · expansion-cohort · GPT-5 mini · expansion-cohort ·
rampstackco-skill-creation-walkthrough · Wider run, three waysNot runNot runNot runhelped · +0.255helped · +0.382
anthropics-skill-creator · Wider run, three waysNot runNot runNot runhelped · +0.255helped · +0.336
rampstackco-skill-creation-walkthrough · First run, with and without the skillhelped · +0.382unclear · +0.127helped · +0.282Not runNot run
yusufkaraaslan-skill-builder · Wider run, three waysNot runNot runNot runhelped · +0.255helped · +0.273
anthropics-skill-creator · First run, with and without the skillhelped · +0.382unclear · +0.127helped · +0.273Not runNot run
openclaw-skill-creator · Wider run, three waysNot runNot runNot runhelped · +0.255unclear · +0.192
Skill authoring · tested in Claude Code
SkillClaude Fable 5 · skill-cohort · Claude Opus 5 · skill-cohort · Claude Fable 5 · panel-v2 · Claude Haiku 4.5 · panel-v2 · Claude Opus 5 · panel-v2 · Claude Sonnet 5 · panel-v2 · Claude Fable 5.1 · fable-5-1 · Claude Fable 5.1 · expansion-cohort · Claude Haiku 4.5 · expansion-cohort · Claude Opus 5 · expansion-cohort · Claude Sonnet 5 · expansion-cohort ·
rampstackco-skill-creation-walkthrough · Wider run, three waysat the ceiling · -0.073at the ceiling · +0.000at the ceiling · -0.073helped · +0.309at the ceiling · +0.000at the ceiling · -0.109at the ceiling · +0.000at the ceiling · -0.145helped · +0.455at the ceiling · +0.000at the ceiling · -0.145
anthropics-skill-creator · Wider run, three waysat the ceiling · +0.000at the ceiling · +0.000at the ceiling · +0.000helped · +0.291at the ceiling · -0.091at the ceiling · -0.073at the ceiling · -0.218at the ceiling · -0.436helped · +0.436at the ceiling · +0.000at the ceiling · +0.000
rampstackco-skill-creation-walkthrough · First run, with and without the skillat the ceiling · -0.073at the ceiling · +0.000at the ceiling · -0.073helped · +0.309at the ceiling · +0.000at the ceiling · -0.109at the ceiling · +0.000at the ceiling · -0.145helped · +0.455at the ceiling · +0.000at the ceiling · -0.145
yusufkaraaslan-skill-builder · Wider run, three waysNot runNot runNot runNot runNot runNot runNot runat the ceiling · +0.000helped · +0.400at the ceiling · +0.000at the ceiling · +0.000
anthropics-skill-creator · First run, with and without the skillat the ceiling · +0.000at the ceiling · +0.000at the ceiling · +0.000helped · +0.291at the ceiling · -0.091at the ceiling · -0.073at the ceiling · -0.218at the ceiling · -0.436helped · +0.436at the ceiling · +0.000at the ceiling · +0.000
openclaw-skill-creator · Wider run, three waysNot runNot runNot runNot runNot runNot runNot runat the ceiling · +0.000helped · +0.345at the ceiling · +0.000at the ceiling · +0.000

Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.

How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in anthropics/skills skills/skill-creator at 3b3fad96af16.

Two candidates name the class. The task class is scored on a single SKILL.md whose frontmatter must parse and whose section order must match SKILL_AUTHORING.md. writing-skills prescribes that document's section order under the heading above. skill-creator prescribes an authoring and evaluation workflow, and the structure it prescribes under its own Report structure heading is an evaluation report rather than the SKILL.md. writing-skills is the closer output format.

Addendum. RECORDED AFTER THE FACT, ON 2026-08-26, AND NOT ACTED ON. This tiebreak was decided when the skill-authoring class was scored against SKILL_AUTHORING.md alone. That document is now Key B, disclosed and unranked, and the ranked key is the open Agent Skills specification. The tiebreak reasoning above therefore turns on a basis that is no longer the ranked one. It is left exactly as it was decided, because a selection record that is rewritten to agree with a later ruling is not a record of what was decided. The selection was NOT re-run against Key A, and whether it would produce the same candidate is an open question stated here rather than assumed.

Second addendum. SECOND ADDENDUM, 2026-08-26. THE TIEBREAK WAS RE-RUN WITH KEY A AS THE OUTPUT-FORMAT BASIS, AND IT SWAPPED. Basis: the selection rule is unchanged, but the output format the class is scored on is now the open Agent Skills specification rather than SKILL_AUTHORING.md. Candidates compared: the same two that qualified, obra/superpowers skills/writing-skills and anthropics/skills skills/skill-creator. Result: skill-creator. It prescribes the specification's own model of the artifact, naming the required and the optional frontmatter fields, the bundled directory layout, the progressive disclosure levels and the line ceiling. writing-skills prescribes a body section order, which is the one part of a SKILL.md the specification explicitly leaves unconstrained, and it disagrees with the specification in three places: it puts the thousand and twenty four character limit on the whole frontmatter rather than on the description, it tells an author the description must not say what the skill does where the specification asks for both, and the example name in its own template is capitalised where the specification allows lowercase only. Run through Key A, writing-skills' own template fails a predicate and skill-creator's passes every applicable one. The original record and the first addendum above are left exactly as written. obra/superpowers skills/writing-skills moves to the rejected pins, where its bytes stay committed and its quoted sentence stays resolvable.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/writing-skills from obra/superpowersb36e0829c6d0d34db5c8aed6yesNames the class in its when-to-use directly.
skills/skill-creator from anthropics/skills3b3fad96af16e51ce07ea7fdyesAlso names the class directly, so the tiebreak applies.
skills/skill-scout from affaan-m/ECCd8409a4b081371b1e8a4aedenoSearches for an existing skill before one is written. It names the moment before authoring, not authoring.
skills/skill-comply from affaan-m/ECCd8409a4b0813e26cbbc30a32noMeasures whether a written skill is obeyed. It is about a skill, not about writing one.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in openclaw/openclaw skills/skill-creator at 61b08f0ebb7a.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/skill-creator from openclaw/openclaw61b08f0ebb7aeea1a8321d20yesNames authoring and reviewing SKILL.md files as such, with no domain narrowing. Pinned.
distribution/claude-plugin/skills/skill-builder from yusufkaraaslan/Skill_Seekersf3972efa33faeea1a8321d20yesNames creating skills in its when-to-use. Its inputs are source materials rather than an authoring brief, which is where it sits relative to the others; under R23 that is ordering and not exclusion. Pinned.
skills/microsoft-skill-creator from github/awesome-copilotc956566a35c32b1e9f91f708noNames the class and then narrows it to one vendor's technologies, and declares a network dependency on the Learn MCP server. Rejected on the same ground as affaan-m/ECC flutter-dart-code-review and wshobson/agents screen-reader-testing: a class-general skill bound to one vendor's surface is not the class. NOT a tiebreak loss.
  • addyosmani/agent-skills skills/using-agent-skills names discovering and invoking skills, not authoring one, the same ground affaan-m/ECC skills/skill-scout was rejected on in the committed record.
  • browser-act/skills browser-act-skill-forge names forging a SKILL.md and conditions it throughout on website exploration and bulk extraction.
  • mattpocock/skills has no skill whose when-to-use names authoring a skill.
  • wshobson/agents has none: its skill-forge-essentials plugin holds ai-debt-detector, session-guard and visual-edit-precision, and none names authoring.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • alirezarezvani/claude-skills has none.
  • Jeffallan/claude-skills has none.
  • EveryInc/compound-engineering-plugin has none.
  • open-gsd/gsd-core has none.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in yusufkaraaslan/Skill_Seekers distribution/claude-plugin/skills/skill-builder at f3972efa33fa.

PINNED UNDER R23, HAVING QUALIFIED AND BEEN DISPLACED BY THE CAP. This skill appears in the recon's skill-authoring qualifying table and was not proposed for a slot, because the class could hold one external skill and openclaw/openclaw skill-creator was the closer match: it names authoring and reviewing SKILL.md files as such, where this one's inputs are documentation, repositories, PDFs and videos, so its subject is building a skill FROM source material rather than authoring one to a brief. That is still the honest comparison and it is no longer a reason for absence.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/skill-creator from openclaw/openclaw61b08f0ebb7a42c7e6285ae8yesNames authoring and reviewing SKILL.md files as such, with no domain narrowing. Pinned.
distribution/claude-plugin/skills/skill-builder from yusufkaraaslan/Skill_Seekersf3972efa33fa42c7e6285ae8yesNames creating skills in its when-to-use. Its inputs are source materials rather than an authoring brief, which is where it sits relative to the others; under R23 that is ordering and not exclusion. Pinned.
skills/microsoft-skill-creator from github/awesome-copilotc956566a35c32b1e9f91f708noNames the class and then narrows it to one vendor's technologies, and declares a network dependency on the Learn MCP server. Rejected on the same ground as affaan-m/ECC flutter-dart-code-review and wshobson/agents screen-reader-testing: a class-general skill bound to one vendor's surface is not the class. NOT a tiebreak loss.
  • addyosmani/agent-skills skills/using-agent-skills names discovering and invoking skills, not authoring one, the same ground affaan-m/ECC skills/skill-scout was rejected on in the committed record.
  • browser-act/skills browser-act-skill-forge names forging a SKILL.md and conditions it throughout on website exploration and bulk extraction.
  • mattpocock/skills has no skill whose when-to-use names authoring a skill.
  • wshobson/agents has none: its skill-forge-essentials plugin holds ai-debt-detector, session-guard and visual-edit-precision, and none names authoring.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • alirezarezvani/claude-skills has none.
  • Jeffallan/claude-skills has none.
  • EveryInc/compound-engineering-plugin has none.
  • open-gsd/gsd-core has none.

Accessibility audit

10 fixture web pages, each seeded with known WCAG defects, 48 in total. The model is asked to list the defects it finds; the score is the share of seeded defects it names. 23 defects are verified in the fixtures by mechanical check; 25 are judgement items the check cannot verify and are listed as such.

4 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.

Every skill tested for accessibility, ranked.

Skill cohort (launch cells)

one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.

  • Helped on someBest result in this test

    skills/accessibility-audit rampstackco/claude-skills*

    Helped on Claude Haiku 4.5; no reliable difference on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.

    Found 71% of the seeded accessibility defects, against 58% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
    • GPT-5 mini · unclear. The model scored 69% without it.

    Adds about 11,700 tokens of context to every request.

  • No reliable difference

    skills/accessibility affaan-m/ECC

    No reliable difference on any model tested. The models scored 58% without it. Measured against no skill loaded.

    Found 56% of the seeded accessibility defects, against 58% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · unclear. The model scored 39% without it.
    • Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
    • GPT-5 mini · unclear. The model scored 71% without it.

    Adds about 2,200 tokens of context to every request.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.

  • No reliable differenceBest result in this test

    skills/accessibility-audit rampstackco/claude-skills*

    No reliable difference on any model tested. The models scored 60% without it. Measured against the same task with a one-line instruction.

    Found 68% of the seeded accessibility defects, against 60% with no skill loaded. Against 60% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
    • GPT-5 mini · unclear. The model scored 57% without it.

    Adds about 11,300 tokens of context to every request.

  • No reliable difference

    plugins/accessibility-compliance/skills/wcag-audit-patterns wshobson/agents

    No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.

    Found 63% of the seeded accessibility defects, against 63% with no skill loaded. Against 58% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
    • GPT-5 mini · unclear. The model scored 64% without it.

    Adds about 4,000 tokens of context to every request.

  • No reliable difference

    engineering-team/a11y-audit/skills/a11y-audit alirezarezvani/claude-skills

    No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction. Based on 60 of 62 runs; the rest returned nothing.

    Found 65% of the seeded accessibility defects, against 64% with no skill loaded. Against 62% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
    • GPT-5 mini · unclear. The model scored 64% without it.

    Adds about 16,400 tokens of context to every request.

  • No reliable difference

    skills/accessibility affaan-m/ECC

    No reliable difference on any model tested. The models scored 55% without it. Measured against the same task with a one-line instruction. Based on 60 of 61 runs; the rest returned nothing.

    Found 57% of the seeded accessibility defects, against 55% with no skill loaded. Against 64% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
    • GPT-5 mini · unclear. The model scored 48% without it.

    Adds about 2,100 tokens of context to every request.

Show the numbers
SkillCohortWithout the skillWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loaded First run, with and without the skillbaseline58%58%0 by definition02026-08-30
One-line instruction instead First run, with and without the skillbaselineinstruction arm: not measured0 by definition02026-08-30
skills/accessibility-audit from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this testrampstackco/claude-skills*58%58% -> 71%+0.13 [-0.00, +0.29]In free drift on 2 of 311,7292026-08-30
skills/accessibility from affaan-m/ECC d8409a4b0813 First run, with and without the skillaffaan-m/ECC58%58% -> 56%-0.02 [-0.10, +0.13]In free drift2,2062026-08-30
No skill loaded Wider run, three waysbaseline61%61%0 by definition02026-09-05
One-line instruction instead Wider run, three waysbaseline61%0 by definition02026-09-05
skills/accessibility-audit from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this testrampstackco/claude-skills*60%60% -> 68%+0.08 [+0.05, +0.10]In free drift11,2952026-09-05
plugins/accessibility-compliance/skills/wcag-audit-patterns from wshobson/agents 38e19c20d2b1 Wider run, three wayswshobson/agents63%63% -> 63%+0.05 [+0.02, +0.09]In free drift3,9672026-09-05
engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills 19392f7a0826 Wider run, three waysalirezarezvani/claude-skills64%64% -> 65%+0.03 [-0.03, +0.09]In free drift16,4072026-09-05
skills/accessibility from affaan-m/ECC d8409a4b0813 Wider run, three waysaffaan-m/ECC55%55% -> 57%-0.06 [-0.13, +0.01]In free drift2,1082026-09-05
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

engineering-team/a11y-audit/skills/a11y-audit Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.763. Instruction only: 0.673. Change +0.130. It could sit between -0.018 and +0.278.

skills/accessibility-audit Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.728. Instruction only: 0.673. Change +0.094. It could sit between -0.068 and +0.257.

skills/accessibility-audit Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.728. Change +0.094. It could sit between -0.068 and +0.257.

skills/accessibility-audit Claude Haiku 4.5 First run, with and without the skill, Deterministic

Holds

No skill: 0.410. With the skill: 0.699. Change +0.289. It could sit between +0.145 and +0.433.

plugins/accessibility-compliance/skills/wcag-audit-patterns Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.690. Instruction only: 0.673. Change +0.057. It could sit between -0.085 and +0.199.

skills/accessibility-audit GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.689. Change -0.003. It could sit between -0.112 and +0.105.

skills/accessibility-audit GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.626. Instruction only: 0.529. Change +0.051. It could sit between -0.074 and +0.175.

skills/accessibility GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.705. With the skill: 0.622. Change -0.083. It could sit between -0.193 and +0.028.

skills/accessibility GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.476. With the skill: 0.608. Instruction only: 0.597. Change +0.133. It could sit between -0.030 and +0.295.

plugins/accessibility-compliance/skills/wcag-audit-patterns GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.635. With the skill: 0.575. Instruction only: 0.488. Change -0.060. It could sit between -0.198 and +0.078.

engineering-team/a11y-audit/skills/a11y-audit GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.641. With the skill: 0.542. Instruction only: 0.568. Change -0.099. It could sit between -0.230 and +0.032.

skills/accessibility Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.538. Instruction only: 0.673. Change -0.095. It could sit between -0.175 and -0.015.

skills/accessibility Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.633. With the skill: 0.538. Change -0.095. It could sit between -0.175 and -0.015.

skills/accessibility Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 0.390. With the skill: 0.521. Change +0.131. It could sit between +0.019 and +0.243.

In Claude Code tested in Claude Code

skills/accessibility-audit Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.875. With the skill: 0.935. Change +0.060. It could sit between -0.024 and +0.144.

engineering-team/a11y-audit/skills/a11y-audit Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.895. With the skill: 0.915. Change +0.020. It could sit between -0.071 and +0.111.

skills/accessibility Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.877. With the skill: 0.915. Change +0.038. It could sit between -0.054 and +0.130.

skills/accessibility Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.800. With the skill: 0.900. Change +0.100. It could sit between -0.096 and +0.296.

skills/accessibility-audit Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.855. With the skill: 0.895. Change +0.040. It could sit between -0.012 and +0.092.

engineering-team/a11y-audit/skills/a11y-audit Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.882. With the skill: 0.890. Change +0.008. It could sit between -0.078 and +0.094.

skills/accessibility Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.857. With the skill: 0.880. Change +0.023. It could sit between -0.047 and +0.094.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.895. With the skill: 0.875. Change -0.020. It could sit between -0.059 and +0.019.

skills/accessibility Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.875. Change -0.060. It could sit between -0.144 and +0.024.

skills/accessibility-audit Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.890. With the skill: 0.870. Change -0.020. It could sit between -0.143 and +0.103.

skills/accessibility Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.842. With the skill: 0.850. Change +0.008. It could sit between -0.078 and +0.094.

skills/accessibility-audit Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.890. With the skill: 0.840. Change -0.050. It could sit between -0.178 and +0.078.

skills/accessibility Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.890. With the skill: 0.837. Change -0.053. It could sit between -0.126 and +0.019.

skills/accessibility-audit Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.910. With the skill: 0.830. Change -0.080. It could sit between -0.200 and +0.040.

skills/accessibility Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.811. With the skill: 0.823. Change +0.012. It could sit between -0.038 and +0.061.

engineering-team/a11y-audit/skills/a11y-audit Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.809. With the skill: 0.807. Change -0.002. It could sit between -0.087 and +0.084.

skills/accessibility-audit Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.838. With the skill: 0.804. Change -0.033. It could sit between -0.111 and +0.044.

skills/accessibility-audit Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.800. With the skill: 0.800. Change +0.000. It could sit between +0.000 and +0.000.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.857. With the skill: 0.790. Change -0.067. It could sit between -0.214 and +0.080.

skills/accessibility Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.838. With the skill: 0.784. Change -0.053. It could sit between -0.164 and +0.058.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.818. With the skill: 0.779. Change -0.038. It could sit between -0.163 and +0.087.

skills/accessibility-audit Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.835. With the skill: 0.777. Change -0.058. It could sit between -0.140 and +0.025.

skills/accessibility-audit Claude Haiku 4.5 Second run, every task twice, Deterministic

Holds

No skill: 0.458. With the skill: 0.690. Change +0.232. It could sit between +0.088 and +0.377.

skills/accessibility-audit Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.455. With the skill: 0.628. Change +0.173. It could sit between +0.030 and +0.317.

engineering-team/a11y-audit/skills/a11y-audit Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.419. With the skill: 0.624. Change +0.205. It could sit between +0.097 and +0.313.

plugins/accessibility-compliance/skills/wcag-audit-patterns Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.477. With the skill: 0.608. Change +0.131. It could sit between +0.021 and +0.240.

skills/accessibility Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.452. With the skill: 0.553. Change +0.101. It could sit between -0.045 and +0.247.

skills/accessibility Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.435. With the skill: 0.513. Change +0.078. It could sit between -0.002 and +0.159.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.

In Claude Code

These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.

Skill cohort (one pass)

one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.

Verdict withheld on this run

skills/accessibility-audit, Claude Fable 5: served model mismatch; 1 usable pair. Kept and not scored into any cell.

skills/accessibility, Claude Fable 5: served model mismatch; 2 usable pairs. Kept and not scored into any cell.

Panel v2 (repeat sampling)

two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.

  • skills/accessibility-audit

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Fable 5 · unclear under repeat sampling
    • Claude Haiku 4.5 · helped under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
  • skills/accessibility

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Fable 5 · unclear under repeat sampling
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling

Fable 5.1, one pass

one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

  • skills/accessibility-audit

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · helped under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · unclear in the version re-test
  • skills/accessibility

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · unclear in the version re-test
  • plugins/accessibility-compliance/skills/wcag-audit-patterns

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • engineering-team/a11y-audit/skills/a11y-audit

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · helped in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
Accessibility audit · tested via API
SkillClaude Haiku 4.5 · skill-cohort · Gemini 3.1 Flash Lite · skill-cohort · GPT-5 mini · skill-cohort · Gemini 3.1 Flash Lite · expansion-cohort · GPT-5 mini · expansion-cohort ·
rampstackco-accessibility-audit · First run, with and without the skillhelped · +0.289unclear · +0.094unclear · -0.003Not runNot run
rampstackco-accessibility-audit · Wider run, three waysNot runNot runNot rununclear · +0.054unclear · +0.097
wshobson-wcag-audit-patterns · Wider run, three waysNot runNot runNot rununclear · +0.017unclear · +0.087
alirezarezvani-a11y-audit · Wider run, three waysNot runNot runNot rununclear · +0.090unclear · -0.026
affaan-m-accessibility · First run, with and without the skillunclear · +0.131unclear · -0.095unclear · -0.083Not runNot run
affaan-m-accessibility · Wider run, three waysNot runNot runNot rununclear · -0.135unclear · +0.012
Accessibility audit · tested in Claude Code
SkillClaude Opus 5 · skill-cohort · Claude Fable 5 · skill-cohort · Claude Fable 5 · panel-v2 · Claude Haiku 4.5 · panel-v2 · Claude Opus 5 · panel-v2 · Claude Sonnet 5 · panel-v2 · Claude Fable 5.1 · fable-5-1 · Claude Fable 5.1 · expansion-cohort · Claude Haiku 4.5 · expansion-cohort · Claude Opus 5 · expansion-cohort · Claude Sonnet 5 · expansion-cohort ·
rampstackco-accessibility-audit · First run, with and without the skillunclear · -0.020Withheldserved model mismatch; 1 usable pairunclear · +0.000helped · +0.232unclear · -0.050unclear · -0.058unclear · +0.060unclear · +0.040unclear · +0.173at the ceiling · -0.080unclear · -0.033
rampstackco-accessibility-audit · Wider run, three waysunclear · -0.020Withheldserved model mismatch; 1 usable pairunclear · +0.000helped · +0.232unclear · -0.050unclear · -0.058unclear · +0.060unclear · +0.040unclear · +0.173at the ceiling · -0.080unclear · -0.033
wshobson-wcag-audit-patterns · Wider run, three waysNot runNot runNot runNot runNot runNot runNot rununclear · -0.020unclear · +0.131unclear · -0.067unclear · -0.038
alirezarezvani-a11y-audit · Wider run, three waysNot runNot runNot runNot runNot runNot runNot rununclear · +0.020helped · +0.205unclear · +0.008unclear · -0.002
affaan-m-accessibility · First run, with and without the skillunclear · -0.053Withheldserved model mismatch; 2 usable pairsunclear · +0.100unclear · +0.101unclear · +0.023unclear · +0.012unclear · +0.008at the ceiling · -0.060unclear · +0.078unclear · +0.038unclear · -0.053
affaan-m-accessibility · Wider run, three waysunclear · -0.053Withheldserved model mismatch; 2 usable pairsunclear · +0.100unclear · +0.101unclear · +0.023unclear · +0.012unclear · +0.008at the ceiling · -0.060unclear · +0.078unclear · +0.038unclear · -0.053

Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.

How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/accessibility at d8409a4b0813.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/accessibility from affaan-m/ECCd8409a4b0813ab86a50717d6yesNames auditing against WCAG in the when-to-use itself.
skills/click-path-audit from affaan-m/ECCd8409a4b0813662fc5d02483noAn audit of behavioural state, not of accessibility. The word audit is shared and the class is not.
  • obra/superpowers has no skill naming accessibility in any when-to-use.
  • anthropics/skills has none either; webapp-testing names browser testing rather than accessibility.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in wshobson/agents plugins/accessibility-compliance/skills/wcag-audit-patterns at 38e19c20d2b1.

TWO SKILLS IN ONE REPOSITORY, AND THE COMPARISON BETWEEN THEM IS NOT A TIEBREAK. wshobson/agents also carries plugins/ui-design/skills/accessibility-compliance, which the recon's table lists as qualifying. It is recorded here as NOT qualifying, and that is a judgment about the class rather than a ranking: its deliverable is a compliant interface, it reaches auditing only as the first of four alternatives, and the artifact this class scores is a findings record of locators against success criteria, which a skill that produces code does not produce. R23 removed the per-class cap, so nothing about that decision turns on there being room for one skill: were it a tiebreak loss it would be pinned alongside this slot like every other qualifying candidate. It is a rejection, its bytes are committed under referencePins.rejected, and its sentence is quoted there.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
plugins/accessibility-compliance/skills/wcag-audit-patterns from wshobson/agents38e19c20d2b17c2619d811feyesThe first clause is the class label. The strongest names-the-class match found anywhere in the pass, for any class. Pinned.
engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills19392f7a08267c2619d811feyesNames auditing accessibility in the first clause of its when-to-use. Pinned.
plugins/ui-design/skills/accessibility-compliance from wshobson/agents38e19c20d2b13ad0fbcb0a69noDOES NOT QUALIFY, AND THIS IS A QUALIFICATION JUDGMENT RATHER THAN A TIEBREAK LOSS. Its deliverable is a compliant interface: it leads with implementing interfaces with inclusive design and assistive-technology support, and reaches auditing only as the first of four alternatives. The class is an audit, and the artifact this class scores is a findings record of locators against success criteria, which a skill that produces code does not produce. R23 removed the cap and did not turn this into an admission: it was never displaced by a better candidate, it was judged not to name the class.
plugins/accessibility-compliance/skills/screen-reader-testing from wshobson/agents38e19c20d2b1591c7ff80f40noOne assistive-technology modality tested directly. Not an audit against a standard.
  • mattpocock/skills has no skill whose when-to-use names accessibility.
  • addyosmani/agent-skills has none.
  • openclaw/openclaw has none.
  • github/awesome-copilot has none.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • Jeffallan/claude-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in alirezarezvani/claude-skills engineering-team/a11y-audit/skills/a11y-audit at 19392f7a0826.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
plugins/accessibility-compliance/skills/wcag-audit-patterns from wshobson/agents38e19c20d2b134f619aeb9ccyesThe first clause is the class label. The strongest names-the-class match found anywhere in the pass, for any class. Pinned.
engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills19392f7a082634f619aeb9ccyesNames auditing accessibility in the first clause of its when-to-use. Pinned.
plugins/ui-design/skills/accessibility-compliance from wshobson/agents38e19c20d2b13ad0fbcb0a69noDOES NOT QUALIFY, AND THIS IS A QUALIFICATION JUDGMENT RATHER THAN A TIEBREAK LOSS. Its deliverable is a compliant interface: it leads with implementing interfaces with inclusive design and assistive-technology support, and reaches auditing only as the first of four alternatives. The class is an audit, and the artifact this class scores is a findings record of locators against success criteria, which a skill that produces code does not produce. R23 removed the cap and did not turn this into an admission: it was never displaced by a better candidate, it was judged not to name the class.
plugins/accessibility-compliance/skills/screen-reader-testing from wshobson/agents38e19c20d2b1591c7ff80f40noOne assistive-technology modality tested directly. Not an audit against a standard.
  • mattpocock/skills has no skill whose when-to-use names accessibility.
  • addyosmani/agent-skills has none.
  • openclaw/openclaw has none.
  • github/awesome-copilot has none.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • Jeffallan/claude-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • yusufkaraaslan/Skill_Seekers has none.

Spec writing

The model is given a feature request and asked for a product spec meeting a stated contract: required sections and testable acceptance criteria. Scored two ways: against the contract, and by a grader comparing the with and without outputs blind. Both readings appear.

7 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.

Every skill tested for product management, ranked.

Skill cohort (launch cells)

one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.

  • Helped on someBest result in this test

    skills/product-capability affaan-m/ECC

    Helped clearly on Claude Haiku 4.5 and Gemini 3.1 Flash Lite; unclear on GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 27 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · unclear. The model scored 30% without it.

    Adds about 1,100 tokens of context to every request.

  • Helped clearly

    skills/pm-spec-writing rampstackco/claude-skills*

    Helped clearly on Claude Haiku 4.5, Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 25 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    Adds about 6,800 tokens of context to every request.

  • Nothing left to measureBest result in this test

    skills/product-capability affaan-m/ECC

    Nothing left to measure. The models scored 98% without the skill. Measured against no skill loaded. Models already did well without a skill, so there was little room to improve.

    Found 95% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · Nothing left to measure. The model scored 100% without the skill.
    • Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
    • GPT-5 mini · Nothing left to measure. The model scored 94% without the skill.

    Adds about 1,100 tokens of context to every request.

  • Nothing left to measure

    skills/pm-spec-writing rampstackco/claude-skills*

    Nothing left to measure. The models scored 98% without the skill. Measured against no skill loaded. Models already did well without a skill, so there was little room to improve.

    Found 94% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · Nothing left to measure. The model scored 99% without the skill.
    • Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
    • GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.

    Adds about 6,800 tokens of context to every request.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.

  • Nothing left to measureBest result in this test

    skills/product-capability affaan-m/ECC

    Nothing left to measure. The models scored 98% without the skill. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.

    Found 98% of the required sections and acceptance criteria, against 98% with no skill loaded. Against 97% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
    • GPT-5 mini · Nothing left to measure. The model scored 97% without the skill.

    Adds about 1,000 tokens of context to every request.

  • Could not be measured

    skills/spec-driven-development addyosmani/agent-skills

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.

    Found 98% of the required sections and acceptance criteria, against 98% with no skill loaded. Against 98% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · Nothing left to measure. The model scored 96% without the skill.

    Adds about 2,800 tokens of context to every request.

  • Could not be measured

    pm-execution/skills/create-prd phuryn/pm-skills

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.

    Found 98% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 98% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.

    Adds about 900 tokens of context to every request.

  • Could not be measured

    engineering/skills/spec-driven-workflow alirezarezvani/claude-skills

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.

    Found 97% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 97% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.

    Adds about 15,000 tokens of context to every request.

  • Could not be measured

    skills/prd github/awesome-copilot

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve. Based on 60 of 62 runs; the rest returned nothing.

    Found 95% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 98% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.

    Adds about 1,100 tokens of context to every request.

  • Nothing left to measure

    skills/pm-spec-writing rampstackco/claude-skills*

    Nothing left to measure. The models scored 99% without the skill. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.

    Found 95% of the required sections and acceptance criteria, against 99% with no skill loaded. Against 98% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
    • GPT-5 mini · Nothing left to measure. The model scored 99% without the skill.

    Adds about 6,700 tokens of context to every request.

Show the numbers
SkillCohortWithout the skillWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loaded Blind comparison of the same task, run twice, with ties counting half, First run, with and without the skillbaseline13%13%0 by definition02026-08-30
One-line instruction instead Blind comparison of the same task, run twice, with ties counting half, First run, with and without the skillbaselineinstruction arm: not applicable to a blind preference0 by definition02026-08-30
skills/product-capability from affaan-m/ECC d8409a4b0813 Blind comparison of the same task, run twice, with ties counting half First run, with and without the skill Best result in this testaffaan-m/ECC10%10% -> 90% baseline shown as the complement+0.80 [+0.40, +1.00]Stable on 2 of 31,0822026-08-30
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Blind comparison of the same task, run twice, with ties counting half First run, with and without the skillrampstackco/claude-skills*17%17% -> 83% baseline shown as the complement+0.67 [+0.60, +0.80]Stable6,8452026-08-30
No skill loaded Deterministic, scored against a committed answer key, Wider run, three waysbaseline98%98%0 by definition02026-09-05
One-line instruction instead Deterministic, scored against a committed answer key, Wider run, three waysbaseline98%0 by definition02026-09-05
skills/product-capability from affaan-m/ECC d8409a4b0813 Deterministic, scored against a committed answer key Wider run, three ways Best result in this testaffaan-m/ECC98% at the ceiling98% -> 98%+0.01 [-0.01, +0.03]In free drift1,0472026-09-05
skills/spec-driven-development from addyosmani/agent-skills d2c37ef6225d Deterministic, scored against a committed answer key Wider run, three waysaddyosmani/agent-skills98% at the ceiling98% -> 98%+0.00 [+0.00, +0.00]Mixed, see page2,7772026-09-05
pm-execution/skills/create-prd from phuryn/pm-skills 18468a95b427 Deterministic, scored against a committed answer key Wider run, three waysphuryn/pm-skills97% at the ceiling97% -> 98%+0.00 [+0.00, +0.00]Mixed, see page9412026-09-05
engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills 19392f7a0826 Deterministic, scored against a committed answer key Wider run, three waysalirezarezvani/claude-skills97% at the ceiling97% -> 97%-0.00 [-0.01, +0.00]Mixed, see page15,0432026-09-05
skills/prd from github/awesome-copilot c956566a35c3 Deterministic, scored against a committed answer key Wider run, three waysgithub/awesome-copilot97% at the ceiling97% -> 95%-0.03 [-0.06, +0.00]Mixed, see page1,1472026-09-05
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Deterministic, scored against a committed answer key Wider run, three waysrampstackco/claude-skills*99% at the ceiling99% -> 95%-0.04 [-0.06, -0.01]In free drift6,6532026-09-05
No skill loaded Deterministic, scored against a committed answer key, First run, with and without the skillbaseline98%98%0 by definition02026-08-30
One-line instruction instead Deterministic, scored against a committed answer key, First run, with and without the skillbaselineinstruction arm: not measured0 by definition02026-08-30
skills/product-capability from affaan-m/ECC d8409a4b0813 Deterministic, scored against a committed answer key First run, with and without the skill Best result in this testaffaan-m/ECC98% at the ceiling98% -> 95%-0.02 [-0.04, +0.00]In free drift1,0822026-08-30
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Deterministic, scored against a committed answer key First run, with and without the skillrampstackco/claude-skills*98% at the ceiling98% -> 94%-0.04 [-0.12, +0.00]In free drift6,8452026-08-30
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

engineering/skills/spec-driven-workflow Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

pm-execution/skills/create-prd Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

skills/prd Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

skills/product-capability Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.000. With the skill: 1.000. Change +1.000. It could sit between +1.000 and +1.000.

skills/product-capability Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.000. With the skill: 1.000. Change +1.000. It could sit between +1.000 and +1.000.

skills/spec-driven-development Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

skills/pm-spec-writing Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Instruction only: 1.000. Change +0.000. It could sit between -0.027 and +0.027.

skills/pm-spec-writing Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Instruction only: 1.000. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Change +0.000. It could sit between -0.027 and +0.027.

skills/pm-spec-writing Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.973. Instruction only: 0.945. Change +0.000. It could sit between -0.038 and +0.038.

skills/product-capability Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.964. Change -0.036. It could sit between -0.065 and -0.007.

pm-execution/skills/create-prd GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.955. Instruction only: 0.955. Change +0.009. It could sit between -0.023 and +0.041.

skills/spec-driven-development GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 0.955. Instruction only: 0.955. Change -0.009. It could sit between -0.051 and +0.032.

engineering/skills/spec-driven-workflow GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.936. Instruction only: 0.945. Change -0.018. It could sit between -0.054 and +0.017.

skills/pm-spec-writing Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.100. With the skill: 0.900. Change +0.800. It could sit between +0.408 and +1.192.

skills/pm-spec-writing GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.900. Instruction only: 0.964. Change -0.091. It could sit between -0.156 and -0.026.

skills/product-capability GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.900. Change -0.036. It could sit between -0.175 and +0.102.

skills/prd GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.891. Instruction only: 0.955. Change -0.064. It could sit between -0.234 and +0.107.

skills/pm-spec-writing GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.836. Change -0.118. It could sit between -0.254 and +0.017.

skills/pm-spec-writing Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/pm-spec-writing GPT-5 mini First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/product-capability GPT-5 mini First run, with and without the skill, Blind comparison of the same task

No measured effect

No skill: 0.300. With the skill: 0.700. Change +0.400. It could sit between -0.199 and +0.999.

In Claude Code tested in Claude Code

skills/pm-spec-writing Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 1.000. Change +0.036. It could sit between +0.007 and +0.065.

pm-execution/skills/create-prd Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.991. Change +0.036. It could sit between +0.007 and +0.065.

skills/prd Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.991. Change +0.018. It could sit between -0.006 and +0.042.

skills/pm-spec-writing Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.982. With the skill: 0.991. Change +0.009. It could sit between -0.023 and +0.041.

pm-execution/skills/create-prd Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 0.982. Change +0.018. It could sit between -0.017 and +0.054.

skills/pm-spec-writing Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.996. With the skill: 0.977. Change -0.018. It could sit between -0.038 and +0.002.

skills/prd Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.973. Change +0.018. It could sit between -0.026 and +0.063.

skills/product-capability Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 0.973. Change +0.009. It could sit between -0.023 and +0.041.

skills/product-capability Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.977. With the skill: 0.968. Change -0.009. It could sit between -0.035 and +0.017.

skills/product-capability Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.964. Change -0.009. It could sit between -0.041 and +0.023.

skills/pm-spec-writing Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.964. Change -0.009. It could sit between -0.051 and +0.032.

engineering/skills/spec-driven-workflow Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.955. Change -0.018. It could sit between -0.076 and +0.040.

skills/product-capability Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.955. Change -0.018. It could sit between -0.054 and +0.017.

pm-execution/skills/create-prd Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.955. Change +0.018. It could sit between -0.026 and +0.063.

skills/pm-spec-writing Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.950. With the skill: 0.950. Change +0.000. It could sit between -0.040 and +0.040.

skills/pm-spec-writing Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.946. Change +0.027. It could sit between +0.004 and +0.051.

skills/prd Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.945. Change -0.009. It could sit between -0.051 and +0.032.

skills/spec-driven-development Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.945. Change +0.000. It could sit between -0.027 and +0.027.

pm-execution/skills/create-prd Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.945. Change +0.036. It could sit between +0.007 and +0.065.

skills/pm-spec-writing Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.945. Change +0.009. It could sit between -0.032 and +0.051.

skills/product-capability Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.945. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.946. With the skill: 0.941. Change -0.004. It could sit between -0.029 and +0.020.

skills/pm-spec-writing Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.936. Change +0.027. It could sit between -0.019 and +0.074.

skills/pm-spec-writing Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.927. With the skill: 0.936. Change +0.009. It could sit between -0.032 and +0.051.

skills/product-capability Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.927. With the skill: 0.936. Change +0.009. It could sit between -0.023 and +0.041.

skills/product-capability Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.936. Change +0.018. It could sit between -0.017 and +0.054.

skills/spec-driven-development Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.982. With the skill: 0.936. Change -0.045. It could sit between -0.075 and -0.016.

skills/pm-spec-writing Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.946. With the skill: 0.927. Change -0.018. It could sit between -0.045 and +0.009.

skills/product-capability Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.968. With the skill: 0.927. Change -0.041. It could sit between -0.066 and -0.016.

skills/product-capability Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.932. With the skill: 0.923. Change -0.009. It could sit between -0.035 and +0.017.

skills/product-capability Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.918. Change +0.000. It could sit between -0.027 and +0.027.

skills/pm-spec-writing Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.918. Change -0.018. It could sit between -0.054 and +0.017.

skills/prd Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.909. Change +0.000. It could sit between +0.000 and +0.000.

skills/spec-driven-development Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.909. Change +0.000. It could sit between +0.000 and +0.000.

skills/spec-driven-development Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.909. Change -0.036. It could sit between -0.065 and -0.007.

engineering/skills/spec-driven-workflow Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.900. Change -0.009. It could sit between -0.041 and +0.023.

engineering/skills/spec-driven-workflow Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.900. Change -0.045. It could sit between -0.085 and -0.006.

engineering/skills/spec-driven-workflow Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.818. Change -0.136. It could sit between -0.166 and -0.107.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.

In Claude Code

These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.

Skill cohort (one pass)

one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.

Panel v2 (repeat sampling)

two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.

  • skills/pm-spec-writing

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
  • skills/product-capability

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling

Fable 5.1, one pass

one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

  • skills/pm-spec-writing

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · unclear in the version re-test
  • skills/product-capability

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · unclear in the version re-test
  • pm-execution/skills/create-prd

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • engineering/skills/spec-driven-workflow

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • skills/spec-driven-development

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • skills/prd

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
Spec writing · tested via API
SkillClaude Haiku 4.5 · skill-cohort · Gemini 3.1 Flash Lite · skill-cohort · GPT-5 mini · skill-cohort · Gemini 3.1 Flash Lite · expansion-cohort · GPT-5 mini · expansion-cohort ·
affaan-m-product-capability · paired, First run, with and without the skillhelped · +1.000helped · +1.000unclear · +0.400Not runNot run
rampstackco-pm-spec-writing · paired, First run, with and without the skillhelped · +0.800helped · +0.600helped · +0.600Not runNot run
affaan-m-product-capability · deterministic, Wider run, three waysNot runNot runNot runat the ceiling · -0.009at the ceiling · +0.027
addyosmani-spec-driven-development · deterministic, Wider run, three waysNot runNot runNot runat the ceiling · +0.000at the ceiling · +0.000
phuryn-create-prd · deterministic, Wider run, three waysNot runNot runNot runat the ceiling · +0.000at the ceiling · +0.000
alirezarezvani-spec-driven-workflow · deterministic, Wider run, three waysNot runNot runNot runat the ceiling · +0.000at the ceiling · -0.009
affaan-m-product-capability · deterministic, First run, with and without the skillat the ceiling · -0.036at the ceiling · +0.000at the ceiling · -0.036Not runNot run
github-prd · deterministic, Wider run, three waysNot runNot runNot runat the ceiling · +0.000at the ceiling · -0.064
rampstackco-pm-spec-writing · deterministic, Wider run, three waysNot runNot runNot runat the ceiling · -0.009at the ceiling · -0.064
rampstackco-pm-spec-writing · deterministic, First run, with and without the skillat the ceiling · +0.000at the ceiling · +0.000at the ceiling · -0.118Not runNot run
Spec writing · tested in Claude Code
SkillClaude Fable 5 · skill-cohort · Claude Opus 5 · skill-cohort · Claude Fable 5 · panel-v2 · Claude Haiku 4.5 · panel-v2 · Claude Opus 5 · panel-v2 · Claude Sonnet 5 · panel-v2 · Claude Fable 5.1 · fable-5-1 · Claude Fable 5.1 · expansion-cohort · Claude Haiku 4.5 · expansion-cohort · Claude Opus 5 · expansion-cohort · Claude Sonnet 5 · expansion-cohort ·
affaan-m-product-capability · paired, First run, with and without the skillNot runNot runNot runNot runNot runNot runNot runNot runNot runNot runNot run
rampstackco-pm-spec-writing · paired, First run, with and without the skillNot runNot runNot runNot runNot runNot runNot runNot runNot runNot runNot run
affaan-m-product-capability · deterministic, Wider run, three waysat the ceiling · +0.009at the ceiling · +0.000at the ceiling · -0.041at the ceiling · -0.009at the ceiling · -0.009at the ceiling · -0.004at the ceiling · -0.018at the ceiling · +0.009at the ceiling · -0.009at the ceiling · +0.018at the ceiling · +0.000
addyosmani-spec-driven-development · deterministic, Wider run, three waysNot runNot runNot runNot runNot runNot runNot runat the ceiling · +0.000at the ceiling · -0.045at the ceiling · +0.000at the ceiling · -0.036
phuryn-create-prd · deterministic, Wider run, three waysNot runNot runNot runNot runNot runNot runNot runat the ceiling · +0.036at the ceiling · +0.018at the ceiling · +0.036at the ceiling · +0.018
alirezarezvani-spec-driven-workflow · deterministic, Wider run, three waysNot runNot runNot runNot runNot runNot runNot runat the ceiling · -0.136at the ceiling · -0.018at the ceiling · -0.009at the ceiling · -0.045
affaan-m-product-capability · deterministic, First run, with and without the skillat the ceiling · +0.009at the ceiling · +0.000at the ceiling · -0.041at the ceiling · -0.009at the ceiling · -0.009at the ceiling · -0.004at the ceiling · -0.018at the ceiling · +0.009at the ceiling · -0.009at the ceiling · +0.018at the ceiling · +0.000
github-prd · deterministic, Wider run, three waysNot runNot runNot runNot runNot runNot runNot runat the ceiling · -0.009at the ceiling · +0.018at the ceiling · +0.000at the ceiling · +0.018
rampstackco-pm-spec-writing · deterministic, Wider run, three waysat the ceiling · -0.018at the ceiling · +0.009at the ceiling · -0.018at the ceiling · -0.018at the ceiling · +0.027at the ceiling · +0.000at the ceiling · +0.009at the ceiling · +0.036at the ceiling · -0.009at the ceiling · +0.027at the ceiling · +0.009
rampstackco-pm-spec-writing · deterministic, First run, with and without the skillat the ceiling · -0.018at the ceiling · +0.009at the ceiling · -0.018at the ceiling · -0.018at the ceiling · +0.027at the ceiling · +0.000at the ceiling · +0.009at the ceiling · +0.036at the ceiling · -0.009at the ceiling · +0.027at the ceiling · +0.009

Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.

How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in affaan-m/ECC skills/product-capability at d8409a4b0813.

Two candidates name the class. The task class is scored on a document with required sections and acceptance criteria, so the closer output format is the one that fixes its sections. product-capability declares a canonical artifact and a section-by-section output format under the heading above. doc-coauthoring prescribes a three-stage collaboration and fixes no sections in the document it produces.

Addendum. RECORDED AFTER THE FACT, ON 2026-08-26, AND NOT ACTED ON. This tiebreak was decided when the skill-authoring class was scored against SKILL_AUTHORING.md alone. That document is now Key B, disclosed and unranked, and the ranked key is the open Agent Skills specification. The tiebreak reasoning above therefore turns on a basis that is no longer the ranked one. It is left exactly as it was decided, because a selection record that is rewritten to agree with a later ruling is not a record of what was decided. The selection was NOT re-run against Key A, and whether it would produce the same candidate is an open question stated here rather than assumed.

Second addendum. SECOND ADDENDUM, 2026-08-26. RE-EXAMINED, NOT RE-RUN, AND THE CANDIDATE IS UNCHANGED. Basis: this tiebreak was never decided on the operator's house standard. The spec-writing class is scored on the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms, which is a contract the two arms are given rather than a convention one of them was taught, and R7 did not move it. Key A is the answer key for the skill-authoring class and has no bearing here. Candidates compared: the same two that qualified, affaan-m/ECC skills/product-capability and anthropics/skills skills/doc-coauthoring. Result: unchanged, product-capability. THE FIRST ADDENDUM ABOVE IS WRONG ON THIS SLOT and is left in place rather than edited. It was applied to both output-format tiebreaks at once and says this one was decided against SKILL_AUTHORING.md, which it was not. Correcting it by rewriting would hide that the error was made; this sentence is the correction.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/product-capability from affaan-m/ECCd8409a4b08133e0052802969yesNames writing a specification from product intent.
skills/doc-coauthoring from anthropics/skills3b3fad96af162e47d78846fayesNames technical specs among several document kinds, so the tiebreak applies.
skills/writing-plans from obra/superpowersb36e0829c6d048508f44bbfdnoTakes a spec as its input. It is downstream of the class, not the class.
skills/product-lens from affaan-m/ECCd8409a4b0813d082be7c3dd9noRules itself out in its own words and hands the class to product-capability.
How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md frontmatter, description, in phuryn/pm-skills pm-execution/skills/create-prd at 18468a95b427.

THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
pm-execution/skills/create-prd from phuryn/pm-skills18468a95b4276f9493aa031byesFixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned.
engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills19392f7a08266f9493aa031byesThe only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned.
skills/spec-driven-development from addyosmani/agent-skillsd2c37ef6225d6f9493aa031byesNames the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits.
skills/prd from github/awesome-copilotc956566a35c36f9493aa031byesLists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23.
skills/spec-miner from Jeffallan/claude-skills882ef55e377d0deb55afcc6cnoIts input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class.
  • mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
  • openclaw/openclaw has no skill whose when-to-use names writing a specification.
  • samber/cc-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • wshobson/agents has none.
  • yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md frontmatter, description, in alirezarezvani/claude-skills engineering/skills/spec-driven-workflow at 19392f7a0826.

THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
pm-execution/skills/create-prd from phuryn/pm-skills18468a95b427119991870509yesFixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned.
engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills19392f7a0826119991870509yesThe only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned.
skills/spec-driven-development from addyosmani/agent-skillsd2c37ef6225d119991870509yesNames the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits.
skills/prd from github/awesome-copilotc956566a35c3119991870509yesLists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23.
skills/spec-miner from Jeffallan/claude-skills882ef55e377d0deb55afcc6cnoIts input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class.
  • mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
  • openclaw/openclaw has no skill whose when-to-use names writing a specification.
  • samber/cc-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • wshobson/agents has none.
  • yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in addyosmani/agent-skills skills/spec-driven-development at d2c37ef6225d.

THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
pm-execution/skills/create-prd from phuryn/pm-skills18468a95b427280779b1914eyesFixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned.
engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills19392f7a0826280779b1914eyesThe only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned.
skills/spec-driven-development from addyosmani/agent-skillsd2c37ef6225d280779b1914eyesNames the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits.
skills/prd from github/awesome-copilotc956566a35c3280779b1914eyesLists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23.
skills/spec-miner from Jeffallan/claude-skills882ef55e377d0deb55afcc6cnoIts input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class.
  • mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
  • openclaw/openclaw has no skill whose when-to-use names writing a specification.
  • samber/cc-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • wshobson/agents has none.
  • yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in github/awesome-copilot skills/prd at c956566a35c3.

THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
pm-execution/skills/create-prd from phuryn/pm-skills18468a95b427fc1caf7b8c27yesFixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned.
engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills19392f7a0826fc1caf7b8c27yesThe only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned.
skills/spec-driven-development from addyosmani/agent-skillsd2c37ef6225dfc1caf7b8c27yesNames the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits.
skills/prd from github/awesome-copilotc956566a35c3fc1caf7b8c27yesLists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23.
skills/spec-miner from Jeffallan/claude-skills882ef55e377d0deb55afcc6cnoIts input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class.
  • mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
  • openclaw/openclaw has no skill whose when-to-use names writing a specification.
  • samber/cc-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • wshobson/agents has none.
  • yusufkaraaslan/Skill_Seekers has none.

On-page audit

10 fixture pages seeded with 43 known on-page SEO issues from a closed list of 19 issue types. The model is asked to list the issues; the score is the share of seeded issues it names. 39 verified mechanically, 4 judgement items listed as such.

3 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.

Every skill tested for SEO, ranked.

Skill cohort (launch cells)

one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.

  • Helped on someBest result in this test

    skills/seo-onpage rampstackco/claude-skills*

    Helped on Gemini 3.1 Flash Lite; no reliable difference on Claude Haiku 4.5 and GPT-5 mini. Measured against no skill loaded.

    Found 56% of the seeded on-page issues, against 39% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · unclear. The model scored 24% without it.
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · unclear. The model scored 45% without it.

    Adds about 6,200 tokens of context to every request.

  • No reliable difference

    skills/seo affaan-m/ECC

    No reliable difference on any model tested. The models scored 39% without it. Measured against no skill loaded.

    Found 44% of the seeded on-page issues, against 39% with no skill loaded. Instruction arm: not measured.

    • Claude Haiku 4.5 · unclear. The model scored 24% without it.
    • Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
    • GPT-5 mini · unclear. The model scored 45% without it.

    Adds about 1,600 tokens of context to every request.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.

  • No reliable differenceBest result in this test

    skills/seo-onpage rampstackco/claude-skills*

    No reliable difference on any model tested. The models scored 53% without it. Measured against the same task with a one-line instruction. Based on 60 of 68 runs; the rest returned nothing.

    Found 63% of the seeded on-page issues, against 53% with no skill loaded. Against 53% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
    • GPT-5 mini · unclear. The model scored 58% without it.

    Adds about 6,000 tokens of context to every request.

  • No reliable difference

    marketing-skill/skills/seo-audit alirezarezvani/claude-skills

    No reliable difference on any model tested. The models scored 47% without it. Measured against the same task with a one-line instruction.

    Found 58% of the seeded on-page issues, against 47% with no skill loaded. Against 56% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
    • GPT-5 mini · unclear. The model scored 46% without it.

    Adds about 5,600 tokens of context to every request.

  • No reliable difference

    skills/seo affaan-m/ECC

    No reliable difference on any model tested. The models scored 51% without it. Measured against the same task with a one-line instruction.

    Found 56% of the seeded on-page issues, against 51% with no skill loaded. Against 56% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
    • GPT-5 mini · unclear. The model scored 54% without it.

    Adds about 1,600 tokens of context to every request.

Show the numbers
SkillCohortWithout the skillWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loaded First run, with and without the skillbaseline39%39%0 by definition02026-08-30
One-line instruction instead First run, with and without the skillbaselineinstruction arm: not measured0 by definition02026-08-30
skills/seo-onpage from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this testrampstackco/claude-skills*39%39% -> 56%+0.17 [+0.04, +0.23]In free drift on 2 of 36,1842026-08-30
skills/seo from affaan-m/ECC d8409a4b0813 First run, with and without the skillaffaan-m/ECC39%39% -> 44%+0.05 [+0.00, +0.12]In free drift1,6372026-08-30
No skill loaded Wider run, three waysbaseline51%51%0 by definition02026-09-05
One-line instruction instead Wider run, three waysbaseline55%0 by definition02026-09-05
skills/seo-onpage from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this testrampstackco/claude-skills*53%53% -> 63%+0.10 [+0.08, +0.13]In free drift5,9752026-09-05
marketing-skill/skills/seo-audit from alirezarezvani/claude-skills 19392f7a0826 Wider run, three waysalirezarezvani/claude-skills47%47% -> 58%+0.02 [-0.04, +0.07]In free drift5,6252026-09-05
skills/seo from affaan-m/ECC d8409a4b0813 Wider run, three waysaffaan-m/ECC51%51% -> 56%-0.00 [-0.06, +0.05]In free drift1,5782026-09-05
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

skills/seo-onpage Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

Holds

No skill: 0.488. With the skill: 0.707. Change +0.219. It could sit between +0.046 and +0.392.

skills/seo-onpage Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.488. With the skill: 0.682. Instruction only: 0.556. Change +0.194. It could sit between +0.032 and +0.356.

marketing-skill/skills/seo-audit Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.488. With the skill: 0.626. Instruction only: 0.556. Change +0.138. It could sit between -0.025 and +0.301.

skills/seo Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.488. With the skill: 0.610. Instruction only: 0.556. Change +0.122. It could sit between -0.054 and +0.297.

skills/seo Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.488. With the skill: 0.610. Change +0.122. It could sit between -0.054 and +0.297.

skills/seo-onpage GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.575. Instruction only: 0.495. Change -0.000. It could sit between -0.200 and +0.199.

marketing-skill/skills/seo-audit GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.455. With the skill: 0.534. Instruction only: 0.570. Change +0.079. It could sit between -0.051 and +0.208.

skills/seo GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.536. With the skill: 0.503. Instruction only: 0.561. Change -0.033. It could sit between -0.140 and +0.073.

skills/seo-onpage GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.450. With the skill: 0.492. Change +0.041. It could sit between -0.135 and +0.218.

skills/seo GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.448. With the skill: 0.480. Change +0.032. It could sit between -0.161 and +0.225.

skills/seo-onpage Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 0.239. With the skill: 0.474. Change +0.235. It could sit between -0.032 and +0.501.

skills/seo Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 0.239. With the skill: 0.239. Change +0.000. It could sit between -0.146 and +0.146.

In Claude Code tested in Claude Code

skills/seo-onpage Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.930. With the skill: 0.983. Change +0.054. It could sit between -0.017 and +0.124.

marketing-skill/skills/seo-audit Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.950. With the skill: 0.950. Change +0.000. It could sit between +0.000 and +0.000.

skills/seo Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.888. With the skill: 0.950. Change +0.062. It could sit between -0.074 and +0.198.

skills/seo Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.983. With the skill: 0.950. Change -0.033. It could sit between -0.099 and +0.032.

skills/seo-onpage Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.983. With the skill: 0.950. Change -0.033. It could sit between -0.099 and +0.032.

skills/seo Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.867. With the skill: 0.940. Change +0.073. It could sit between -0.015 and +0.162.

skills/seo-onpage Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.894. With the skill: 0.938. Change +0.043. It could sit between -0.000 and +0.087.

skills/seo Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.942. With the skill: 0.933. Change -0.008. It could sit between -0.094 and +0.078.

skills/seo-onpage Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.905. With the skill: 0.925. Change +0.020. It could sit between -0.064 and +0.105.

marketing-skill/skills/seo-audit Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.867. With the skill: 0.833. Change -0.033. It could sit between -0.143 and +0.076.

skills/seo-onpage Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.892. With the skill: 0.833. Change -0.058. It could sit between -0.187 and +0.071.

skills/seo-onpage Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.879. With the skill: 0.833. Change -0.046. It could sit between -0.113 and +0.021.

skills/seo Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.892. With the skill: 0.832. Change -0.060. It could sit between -0.161 and +0.040.

skills/seo Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.867. With the skill: 0.805. Change -0.062. It could sit between -0.143 and +0.019.

skills/seo-onpage Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.833. With the skill: 0.792. Change -0.042. It could sit between -0.178 and +0.094.

skills/seo Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.867. With the skill: 0.788. Change -0.079. It could sit between -0.161 and +0.004.

skills/seo Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.601. With the skill: 0.665. Change +0.064. It could sit between -0.108 and +0.236.

marketing-skill/skills/seo-audit Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.556. With the skill: 0.654. Change +0.097. It could sit between -0.057 and +0.251.

skills/seo Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.594. With the skill: 0.647. Change +0.053. It could sit between -0.029 and +0.136.

skills/seo-onpage Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.586. With the skill: 0.637. Change +0.051. It could sit between -0.039 and +0.141.

skills/seo-onpage Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.554. With the skill: 0.629. Change +0.075. It could sit between -0.030 and +0.180.

skills/seo-onpage Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.435. With the skill: 0.411. Change -0.023. It could sit between -0.214 and +0.168.

skills/seo-onpage Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.248. With the skill: 0.408. Change +0.161. It could sit between +0.035 and +0.286.

marketing-skill/skills/seo-audit Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.368. With the skill: 0.389. Change +0.021. It could sit between -0.302 and +0.345.

skills/seo Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.343. With the skill: 0.363. Change +0.020. It could sit between -0.203 and +0.242.

skills/seo Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.214. With the skill: 0.323. Change +0.108. It could sit between -0.040 and +0.257.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.

In Claude Code

These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.

Skill cohort (one pass)

one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.

  • skills/seo-onpage

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · unclear, and both runs agree
  • skills/seo

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · unclear, and both runs agree

Panel v2 (repeat sampling)

two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.

  • skills/seo-onpage

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
  • skills/seo

    • Claude Fable 5 · unclear, and both runs agree
    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling

Fable 5.1, one pass

one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

  • skills/seo-onpage

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · unclear in the version re-test
  • skills/seo

    • Claude Opus 5 · unclear, and both runs agree
    • Claude Haiku 4.5 · unclear under repeat sampling
    • Claude Sonnet 5 · unclear under repeat sampling
    • Claude Fable 5.1 · unclear in the version re-test
  • marketing-skill/skills/seo-audit

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
On-page audit · tested via API
SkillClaude Haiku 4.5 · skill-cohort · Gemini 3.1 Flash Lite · skill-cohort · GPT-5 mini · skill-cohort · Gemini 3.1 Flash Lite · expansion-cohort · GPT-5 mini · expansion-cohort ·
rampstackco-seo-onpage · First run, with and without the skillunclear · +0.235helped · +0.219unclear · +0.041Not runNot run
rampstackco-seo-onpage · Wider run, three waysNot runNot runNot rununclear · +0.126unclear · +0.081
affaan-m-seo · First run, with and without the skillunclear · +0.000unclear · +0.122unclear · +0.032Not runNot run
alirezarezvani-seo-audit · Wider run, three waysNot runNot runNot rununclear · +0.070unclear · -0.036
affaan-m-seo · Wider run, three waysNot runNot runNot rununclear · +0.053unclear · -0.058
On-page audit · tested in Claude Code
SkillClaude Fable 5 · skill-cohort · Claude Opus 5 · skill-cohort · Claude Fable 5 · panel-v2 · Claude Haiku 4.5 · panel-v2 · Claude Opus 5 · panel-v2 · Claude Sonnet 5 · panel-v2 · Claude Fable 5.1 · fable-5-1 · Claude Fable 5.1 · expansion-cohort · Claude Haiku 4.5 · expansion-cohort · Claude Opus 5 · expansion-cohort · Claude Sonnet 5 · expansion-cohort ·
rampstackco-seo-onpage · First run, with and without the skillat the ceiling · +0.020unclear · -0.058unclear · +0.043unclear · +0.161unclear · -0.046unclear · +0.051at the ceiling · -0.033at the ceiling · +0.054unclear · -0.023unclear · -0.042unclear · +0.075
rampstackco-seo-onpage · Wider run, three waysat the ceiling · +0.020unclear · -0.058unclear · +0.043unclear · +0.161unclear · -0.046unclear · +0.051at the ceiling · -0.033at the ceiling · +0.054unclear · -0.023unclear · -0.042unclear · +0.075
affaan-m-seo · First run, with and without the skillat the ceiling · -0.008unclear · -0.062unclear · +0.073unclear · +0.108unclear · -0.060unclear · +0.053at the ceiling · -0.033unclear · +0.062unclear · +0.020unclear · -0.079unclear · +0.064
alirezarezvani-seo-audit · Wider run, three waysNot runNot runNot runNot runNot runNot runNot runat the ceiling · +0.000unclear · +0.021unclear · -0.033unclear · +0.097
affaan-m-seo · Wider run, three waysat the ceiling · -0.008unclear · -0.062unclear · +0.073unclear · +0.108unclear · -0.060unclear · +0.053at the ceiling · -0.033unclear · +0.062unclear · +0.020unclear · -0.079unclear · +0.064

Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.

How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/seo at d8409a4b0813.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/seo from affaan-m/ECCd8409a4b08139a655a52cfd9yesNames auditing and on-page optimization in the same sentence.
  • obra/superpowers has no skill naming search or on-page work in any when-to-use.
  • anthropics/skills has none either.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in alirezarezvani/claude-skills marketing-skill/skills/seo-audit at 19392f7a0826.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
marketing-skill/skills/seo-audit from alirezarezvani/claude-skills19392f7a0826979f942727c2yesNames the class label verbatim: "on-page SEO" and "audit" in the same sentence. The most literal match to a class name anywhere in the pass. Pinned.
marketing-skill/skills/programmatic-seo from alirezarezvani/claude-skills19392f7a08260851f7c8ab34noRules itself out in its own words and hands the class to seo-audit, the same shape as affaan-m/ECC skills/product-lens in the committed record.
marketing-skill/skills/local-seo-manager from alirezarezvani/claude-skills19392f7a08266d93fc13f7e0noLocal listings and map presence. Names a channel, not an on-page audit, and hands national SEO to seo-audit in its own words.
  • mattpocock/skills has no skill whose when-to-use names SEO or on-page work.
  • addyosmani/agent-skills has none.
  • openclaw/openclaw has none.
  • wshobson/agents has none.
  • github/awesome-copilot has none.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • Jeffallan/claude-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • yusufkaraaslan/Skill_Seekers has none.

Voice

The model is given a brand brief and asked for a short piece of copy. There is no answer key; a grader sees the with and without outputs, blind and in random order, and picks one. The score is how often the skill's output was preferred.

3 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.

Every skill tested for brand voice, ranked.

Skill cohort (launch cells)

one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.

  • Helped clearlyBest result in this test

    skills/brand-voice affaan-m/ECC

    Helped clearly on Claude Haiku 4.5, Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 24 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    Adds about 1,200 tokens of context to every request.

  • Helped on some

    skills/brand-voice rampstackco/claude-skills*

    Helped on Gemini 3.1 Flash Lite; no reliable difference on Claude Haiku 4.5 and GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 21 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · unclear. The model scored 30% without it.
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · unclear. The model scored 40% without it.

    Adds about 5,700 tokens of context to every request.

Show the numbers
SkillCohortWithout the skillWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loadedbaseline25%25%0 by definition02026-08-30
One-line instruction insteadbaselineinstruction arm: not applicable to a blind preference0 by definition02026-08-30
skills/brand-voice from affaan-m/ECC d8409a4b0813 Best result in this testaffaan-m/ECC20%20% -> 80% baseline shown as the complement+0.60 [+0.60, +0.60]Stable1,1702026-08-30
skills/brand-voice from rampstackco/claude-skills 047924252254rampstackco/claude-skills*30%30% -> 70% baseline shown as the complement+0.40 [+0.20, +0.60]In free drift on 2 of 35,6742026-08-30
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded

Via API tested via API

skills/brand-voice Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice GPT-5 mini First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

No measured effect

No skill: 0.300. With the skill: 0.700. Change +0.400. It could sit between -0.199 and +0.999.

skills/brand-voice GPT-5 mini First run, with and without the skill, Blind comparison of the same task

No measured effect

No skill: 0.400. With the skill: 0.600. Change +0.200. It could sit between -0.440 and +0.840.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.

Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.

How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/brand-voice at d8409a4b0813.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/brand-voice from affaan-m/ECCd8409a4b0813eae455eed766yesNames writing voice and consistency of it.
skills/brand-guidelines from anthropics/skills3b3fad96af161120b3769e29noVisual identity, colors and type. Brand is shared and voice is not.
  • obra/superpowers has no skill naming writing voice in any when-to-use.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in samber/cc-skills skills/copywriting-tone-of-voice-creator at 62bbac4f2f0c.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/copywriting-tone-of-voice-creator from samber/cc-skills62bbac4f2f0cf4753aa281d4yesNames defining brand voice in its when-to-use. It declares a fixed artifact, TONE.md, and a fixed section inventory, which is a closer output-format fit than the incumbent's own description states. Pinned.
marketing-skill/skills/content-humanizer from alirezarezvani/claude-skills19392f7a0826394c104c2199noThe class named is removing AI tells from finished copy, and the skill excludes content creation in its own words. "Voice" appears once, inside a trigger list. Removing a generic register is not building or holding a specific one.
skills/finnish-humanizer from github/awesome-copilotc956566a35c3a21c9d7e4fb7noThe same shape as content-humanizer, narrowed further to one natural language. QUOTATION CORRECTED AGAINST THE BYTES: recon section 2.5 ends this description at a full stop where the committed bytes carry a comma and continue for two further sentences, so the report quotes a truncation. It is recorded here as the artifact writes it.
  • mattpocock/skills has no skill whose when-to-use names writing voice.
  • addyosmani/agent-skills has none.
  • openclaw/openclaw has none.
  • wshobson/agents has none.
  • phuryn/pm-skills has none.
  • Jeffallan/claude-skills has none.
  • open-gsd/gsd-core has none.
  • EveryInc/compound-engineering-plugin has none.
  • yusufkaraaslan/Skill_Seekers has none.

Code review

10 fixture source files in 4 languages, each seeded with known bugs, security issues and correctness problems, 54 in total, drawn from a closed list of 18 defect types that both arms are given. The model is asked to name the symbol each defect sits on, never a line number; the score is the share of seeded defects it names. 50 are verified in the fixtures by mechanical check, 4 are judgement items listed as such, and 30 pieces of correct code that read as suspicious are recorded as traps and are deliberately not in the key.

8 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.

Every skill tested for code review, ranked.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.

  • No reliable differenceBest result in this test

    plugins/developer-essentials/skills/code-review-excellence wshobson/agents

    No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction.

    Found 63% of the seeded code defects, against 64% with no skill loaded. Against 62% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
    • GPT-5 mini · unclear. The model scored 71% without it.

    Adds about 3,900 tokens of context to every request.

  • No reliable difference

    skills/code-reviewer Jeffallan/claude-skills

    No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.

    Found 64% of the seeded code defects, against 63% with no skill loaded. Against 64% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
    • GPT-5 mini · unclear. The model scored 69% without it.

    Adds about 8,300 tokens of context to every request.

  • No reliable difference

    engineering/skills/pr-review-expert alirezarezvani/claude-skills

    No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.

    Found 60% of the seeded code defects, against 63% with no skill loaded. Against 65% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
    • GPT-5 mini · unclear. The model scored 69% without it.

    Adds about 4,300 tokens of context to every request.

  • No reliable difference

    skills/code-review-and-quality addyosmani/agent-skills

    No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction.

    Found 61% of the seeded code defects, against 64% with no skill loaded. Against 66% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
    • GPT-5 mini · unclear. The model scored 71% without it.

    Adds about 5,200 tokens of context to every request.

  • No reliable difference

    skills/gsd-code-review open-gsd/gsd-core

    No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.

    Found 59% of the seeded code defects, against 63% with no skill loaded. Against 64% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
    • GPT-5 mini · unclear. The model scored 68% without it.

    Adds about 1,400 tokens of context to every request.

  • No reliable difference

    skills/code-review-web rampstackco/claude-skills*

    No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction.

    Found 58% of the seeded code defects, against 64% with no skill loaded. Against 65% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
    • GPT-5 mini · unclear. The model scored 71% without it.

    Adds about 8,200 tokens of context to every request.

  • No reliable difference

    skills/engineering/code-review mattpocock/skills

    No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.

    Found 57% of the seeded code defects, against 63% with no skill loaded. Against 69% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
    • GPT-5 mini · unclear. The model scored 69% without it.

    Adds about 2,300 tokens of context to every request.

  • Made worse

    skills/ce-code-review EveryInc/compound-engineering-plugin

    Made results worse on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction.

    Found 34% of the seeded code defects, against 64% with no skill loaded. Against 64% with a one-line instruction instead.

    • Gemini 3.1 Flash Lite · worse
    • GPT-5 mini · unclear. The model scored 71% without it.

    Adds about 73,700 tokens of context to every request.

Show the numbers
SkillCohortWithout the skillWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loadedbaseline64%64%0 by definition02026-09-05
One-line instruction insteadbaseline65%0 by definition02026-09-05
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents 38e19c20d2b1 Best result in this testwshobson/agents64%64% -> 63%+0.01 [+0.01, +0.02]In free drift3,9382026-09-05
skills/code-reviewer from Jeffallan/claude-skills 882ef55e377dJeffallan/claude-skills63%63% -> 64%-0.00 [-0.05, +0.05]In free drift8,2902026-09-05
engineering/skills/pr-review-expert from alirezarezvani/claude-skills 19392f7a0826alirezarezvani/claude-skills63%63% -> 60%-0.05 [-0.08, -0.01]In free drift4,2602026-09-05
skills/code-review-and-quality from addyosmani/agent-skills d2c37ef6225daddyosmani/agent-skills64%64% -> 61%-0.05 [-0.05, -0.04]In free drift5,2182026-09-05
skills/gsd-code-review from open-gsd/gsd-core 6beaa66b2587open-gsd/gsd-core63%63% -> 59%-0.05 [-0.07, -0.03]In free drift1,4202026-09-05
skills/code-review-web from rampstackco/claude-skills a67dd34c609frampstackco/claude-skills*64%64% -> 58%-0.07 [-0.14, -0.01]In free drift8,2062026-09-05
skills/engineering/code-review from mattpocock/skills 6654f6b60cd9mattpocock/skills63%63% -> 57%-0.12 [-0.14, -0.10]In free drift2,2782026-09-05
skills/ce-code-review from EveryInc/compound-engineering-plugin c9c10f8c7541EveryInc/compound-engineering-plugin64%64% -> 34%-0.30 [-0.59, -0.02]Mixed, see page73,7152026-09-05
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

skills/code-reviewer GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.738. Instruction only: 0.690. Change +0.045. It could sit between -0.032 and +0.122.

skills/code-review-web GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.702. Instruction only: 0.710. Change -0.008. It could sit between -0.100 and +0.083.

skills/code-review-and-quality GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.683. Instruction only: 0.727. Change -0.027. It could sit between -0.133 and +0.079.

skills/ce-code-review GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.677. Instruction only: 0.693. Change -0.033. It could sit between -0.077 and +0.010.

skills/gsd-code-review GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.677. With the skill: 0.660. Instruction only: 0.693. Change -0.017. It could sit between -0.124 and +0.091.

plugins/developer-essentials/skills/code-review-excellence GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.658. Instruction only: 0.643. Change -0.052. It could sit between -0.221 and +0.117.

skills/engineering/code-review GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.657. Instruction only: 0.792. Change -0.037. It could sit between -0.085 and +0.011.

engineering/skills/pr-review-expert GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.628. Instruction only: 0.707. Change -0.065. It could sit between -0.171 and +0.041.

plugins/developer-essentials/skills/code-review-excellence Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.602. Instruction only: 0.592. Change +0.027. It could sit between -0.076 and +0.130.

engineering/skills/pr-review-expert Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.578. Instruction only: 0.592. Change +0.003. It could sit between -0.069 and +0.076.

skills/code-review-and-quality Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.542. Instruction only: 0.592. Change -0.033. It could sit between -0.077 and +0.010.

skills/code-reviewer Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.538. Instruction only: 0.592. Change -0.037. It could sit between -0.127 and +0.053.

skills/gsd-code-review Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.522. Instruction only: 0.592. Change -0.053. It could sit between -0.126 and +0.019.

skills/engineering/code-review Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.492. Instruction only: 0.592. Change -0.083. It could sit between -0.189 and +0.022.

skills/code-review-web Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.453. Instruction only: 0.592. Change -0.122. It could sit between -0.239 and -0.005.

skills/ce-code-review Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Scored worse

No skill: 0.575. With the skill: 0.000. Instruction only: 0.592. Change -0.575. It could sit between -0.749 and -0.401.

In Claude Code tested in Claude Code

skills/code-review-and-quality Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.955. Change +0.020. It could sit between -0.019 and +0.059.

plugins/developer-essentials/skills/code-review-excellence Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.935. Change +0.020. It could sit between -0.019 and +0.059.

skills/ce-code-review Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-review-web Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-reviewer Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-reviewer Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.935. Change +0.020. It could sit between -0.019 and +0.059.

skills/engineering/code-review Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.935. Change +0.017. It could sit between -0.016 and +0.049.

skills/engineering/code-review Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.935. Change +0.020. It could sit between -0.019 and +0.059.

skills/gsd-code-review Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/gsd-code-review Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/ce-code-review Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.927. Change +0.008. It could sit between -0.054 and +0.070.

engineering/skills/pr-review-expert Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.918. Change -0.017. It could sit between -0.049 and +0.016.

skills/code-review-and-quality Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.918. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-review-web Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.902. With the skill: 0.902. Change +0.000. It could sit between +0.000 and +0.000.

engineering/skills/pr-review-expert Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.898. Change -0.017. It could sit between -0.049 and +0.016.

plugins/developer-essentials/skills/code-review-excellence Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.862. Change -0.057. It could sit between -0.113 and +0.000.

skills/code-review-web Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.793. With the skill: 0.833. Change +0.040. It could sit between -0.012 and +0.092.

skills/code-reviewer Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.777. With the skill: 0.773. Change -0.003. It could sit between -0.057 and +0.050.

skills/gsd-code-review Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.760. With the skill: 0.763. Change +0.003. It could sit between -0.050 and +0.057.

plugins/developer-essentials/skills/code-review-excellence Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.777. With the skill: 0.760. Change -0.017. It could sit between -0.084 and +0.050.

skills/engineering/code-review Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.760. With the skill: 0.760. Change +0.000. It could sit between -0.049 and +0.049.

skills/code-review-and-quality Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.770. With the skill: 0.757. Change -0.013. It could sit between -0.093 and +0.067.

engineering/skills/pr-review-expert Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.723. With the skill: 0.740. Change +0.017. It could sit between -0.066 and +0.099.

skills/ce-code-review Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.777. With the skill: 0.710. Change -0.067. It could sit between -0.160 and +0.026.

skills/gsd-code-review Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.535. With the skill: 0.567. Change +0.032. It could sit between -0.104 and +0.167.

plugins/developer-essentials/skills/code-review-excellence Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.482. With the skill: 0.553. Change +0.072. It could sit between -0.133 and +0.276.

skills/code-review-web Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.498. With the skill: 0.517. Change +0.018. It could sit between -0.169 and +0.206.

skills/engineering/code-review Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.495. With the skill: 0.507. Change +0.012. It could sit between -0.077 and +0.100.

skills/code-reviewer Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.498. With the skill: 0.503. Change +0.005. It could sit between -0.096 and +0.106.

skills/code-review-and-quality Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.490. With the skill: 0.490. Change -0.000. It could sit between -0.103 and +0.103.

skills/ce-code-review Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.578. With the skill: 0.437. Change -0.142. It could sit between -0.319 and +0.036.

engineering/skills/pr-review-expert Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.488. With the skill: 0.437. Change -0.052. It could sit between -0.317 and +0.214.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.

In Claude Code

These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

  • skills/code-review-web

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • skills/engineering/code-review

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • engineering/skills/pr-review-expert

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • skills/code-review-and-quality

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • skills/code-reviewer

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • skills/ce-code-review

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • plugins/developer-essentials/skills/code-review-excellence

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort
  • skills/gsd-code-review

    • Claude Fable 5.1 · unclear in the three-arm cohort
    • Claude Haiku 4.5 · unclear in the three-arm cohort
    • Claude Opus 5 · unclear in the three-arm cohort
    • Claude Sonnet 5 · unclear in the three-arm cohort

Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.

How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section 3 "Identify the standards sources", in mattpocock/skills skills/engineering/code-review at 6654f6b60cd9.

THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart. THE CANDIDATE LIST HOLDS THE SEVEN THAT QUALIFIED, AND THE NON-QUALIFYING SKILLS ARE IN `emptyRepos` RATHER THAN AS CANDIDATES. Every candidate in this record carries committed bytes, which is what makes the losing side of a decision checkable; a candidate whose bytes cannot be committed would be a quoted sentence with nothing behind it. CherryHQ/cherry-studio .agents/skills/gh-pr-review names the class and is NOT recorded as a candidate for exactly that reason: the repository is AGPL-3.0 at its root, pinning its bytes into this repository would impose AGPL obligations on the surrounding work, and recon section 5.1 bars it. It is named here instead, with its reason, so the exclusion is auditable rather than silent.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/engineering/code-review from mattpocock/skills6654f6b60cd9dcded7969b1dyesNames the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned.
engineering/skills/pr-review-expert from alirezarezvani/claude-skills19392f7a0826dcded7969b1dyesNames the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned.
skills/code-review-and-quality from addyosmani/agent-skillsd2c37ef6225ddcded7969b1dyesNames the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned.
skills/code-reviewer from Jeffallan/claude-skills882ef55e377ddcded7969b1dyesNames the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned.
skills/ce-code-review from EveryInc/compound-engineering-pluginc9c10f8c7541dcded7969b1dyesNames the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned.
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents38e19c20d2b1dcded7969b1dyesNames the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned.
skills/gsd-code-review from open-gsd/gsd-core6beaa66b2587dcded7969b1dyesNames the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned.
  • openclaw/openclaw has no skill whose when-to-use names reviewing code.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • yusufkaraaslan/Skill_Seekers has none.
  • obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
  • affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
  • anthropics/skills has no skill whose when-to-use names reviewing code.
  • github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in alirezarezvani/claude-skills engineering/skills/pr-review-expert at 19392f7a0826.

THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/engineering/code-review from mattpocock/skills6654f6b60cd90b3ac983ef14yesNames the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned.
engineering/skills/pr-review-expert from alirezarezvani/claude-skills19392f7a08260b3ac983ef14yesNames the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned.
skills/code-review-and-quality from addyosmani/agent-skillsd2c37ef6225d0b3ac983ef14yesNames the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned.
skills/code-reviewer from Jeffallan/claude-skills882ef55e377d0b3ac983ef14yesNames the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned.
skills/ce-code-review from EveryInc/compound-engineering-pluginc9c10f8c75410b3ac983ef14yesNames the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned.
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents38e19c20d2b10b3ac983ef14yesNames the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned.
skills/gsd-code-review from open-gsd/gsd-core6beaa66b25870b3ac983ef14yesNames the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned.
  • openclaw/openclaw has no skill whose when-to-use names reviewing code.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • yusufkaraaslan/Skill_Seekers has none.
  • obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
  • affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
  • anthropics/skills has no skill whose when-to-use names reviewing code.
  • github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in addyosmani/agent-skills skills/code-review-and-quality at d2c37ef6225d.

THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/engineering/code-review from mattpocock/skills6654f6b60cd92e6a37430eacyesNames the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned.
engineering/skills/pr-review-expert from alirezarezvani/claude-skills19392f7a08262e6a37430eacyesNames the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned.
skills/code-review-and-quality from addyosmani/agent-skillsd2c37ef6225d2e6a37430eacyesNames the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned.
skills/code-reviewer from Jeffallan/claude-skills882ef55e377d2e6a37430eacyesNames the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned.
skills/ce-code-review from EveryInc/compound-engineering-pluginc9c10f8c75412e6a37430eacyesNames the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned.
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents38e19c20d2b12e6a37430eacyesNames the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned.
skills/gsd-code-review from open-gsd/gsd-core6beaa66b25872e6a37430eacyesNames the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned.
  • openclaw/openclaw has no skill whose when-to-use names reviewing code.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • yusufkaraaslan/Skill_Seekers has none.
  • obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
  • affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
  • anthropics/skills has no skill whose when-to-use names reviewing code.
  • github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in Jeffallan/claude-skills skills/code-reviewer at 882ef55e377d.

THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/engineering/code-review from mattpocock/skills6654f6b60cd9f0f24eedceeeyesNames the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned.
engineering/skills/pr-review-expert from alirezarezvani/claude-skills19392f7a0826f0f24eedceeeyesNames the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned.
skills/code-review-and-quality from addyosmani/agent-skillsd2c37ef6225df0f24eedceeeyesNames the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned.
skills/code-reviewer from Jeffallan/claude-skills882ef55e377df0f24eedceeeyesNames the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned.
skills/ce-code-review from EveryInc/compound-engineering-pluginc9c10f8c7541f0f24eedceeeyesNames the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned.
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents38e19c20d2b1f0f24eedceeeyesNames the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned.
skills/gsd-code-review from open-gsd/gsd-core6beaa66b2587f0f24eedceeeyesNames the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned.
  • openclaw/openclaw has no skill whose when-to-use names reviewing code.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • yusufkaraaslan/Skill_Seekers has none.
  • obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
  • affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
  • anthropics/skills has no skill whose when-to-use names reviewing code.
  • github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in EveryInc/compound-engineering-plugin skills/ce-code-review at c9c10f8c7541.

THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/engineering/code-review from mattpocock/skills6654f6b60cd9bd5c9fac33fbyesNames the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned.
engineering/skills/pr-review-expert from alirezarezvani/claude-skills19392f7a0826bd5c9fac33fbyesNames the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned.
skills/code-review-and-quality from addyosmani/agent-skillsd2c37ef6225dbd5c9fac33fbyesNames the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned.
skills/code-reviewer from Jeffallan/claude-skills882ef55e377dbd5c9fac33fbyesNames the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned.
skills/ce-code-review from EveryInc/compound-engineering-pluginc9c10f8c7541bd5c9fac33fbyesNames the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned.
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents38e19c20d2b1bd5c9fac33fbyesNames the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned.
skills/gsd-code-review from open-gsd/gsd-core6beaa66b2587bd5c9fac33fbyesNames the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned.
  • openclaw/openclaw has no skill whose when-to-use names reviewing code.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • yusufkaraaslan/Skill_Seekers has none.
  • obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
  • affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
  • anthropics/skills has no skill whose when-to-use names reviewing code.
  • github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in wshobson/agents plugins/developer-essentials/skills/code-review-excellence at 38e19c20d2b1.

THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/engineering/code-review from mattpocock/skills6654f6b60cd9d796f1a58081yesNames the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned.
engineering/skills/pr-review-expert from alirezarezvani/claude-skills19392f7a0826d796f1a58081yesNames the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned.
skills/code-review-and-quality from addyosmani/agent-skillsd2c37ef6225dd796f1a58081yesNames the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned.
skills/code-reviewer from Jeffallan/claude-skills882ef55e377dd796f1a58081yesNames the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned.
skills/ce-code-review from EveryInc/compound-engineering-pluginc9c10f8c7541d796f1a58081yesNames the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned.
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents38e19c20d2b1d796f1a58081yesNames the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned.
skills/gsd-code-review from open-gsd/gsd-core6beaa66b2587d796f1a58081yesNames the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned.
  • openclaw/openclaw has no skill whose when-to-use names reviewing code.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • yusufkaraaslan/Skill_Seekers has none.
  • obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
  • affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
  • anthropics/skills has no skill whose when-to-use names reviewing code.
  • github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in open-gsd/gsd-core skills/gsd-code-review at 6beaa66b2587.

THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/engineering/code-review from mattpocock/skills6654f6b60cd99d5072bfb96ayesNames the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned.
engineering/skills/pr-review-expert from alirezarezvani/claude-skills19392f7a08269d5072bfb96ayesNames the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned.
skills/code-review-and-quality from addyosmani/agent-skillsd2c37ef6225d9d5072bfb96ayesNames the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned.
skills/code-reviewer from Jeffallan/claude-skills882ef55e377d9d5072bfb96ayesNames the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned.
skills/ce-code-review from EveryInc/compound-engineering-pluginc9c10f8c75419d5072bfb96ayesNames the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned.
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents38e19c20d2b19d5072bfb96ayesNames the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned.
skills/gsd-code-review from open-gsd/gsd-core6beaa66b25879d5072bfb96ayesNames the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned.
  • openclaw/openclaw has no skill whose when-to-use names reviewing code.
  • phuryn/pm-skills has none.
  • samber/cc-skills has none.
  • yusufkaraaslan/Skill_Seekers has none.
  • obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
  • affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
  • anthropics/skills has no skill whose when-to-use names reviewing code.
  • github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.

Test writing

10 small modules in 2 languages, pure functions and small classes with no input or output of their own, each with a committed list of behaviours it is claimed to obey, 50 in total. The model is asked to write tests for the module and is never shown the list. A behaviour counts as covered only when the model's tests pass against the correct module and fail against a committed broken copy of it that changes exactly that one behaviour; the score is the share of behaviours covered. A test file that fails against the correct module scores zero for that module and the reason is recorded. 20 tests that read as though they cover a behaviour and would catch nothing are committed as traps and are deliberately not in the key.

0 repos have a qualifying skill in this class; 3 declared empty; 12 not yet verified. Which repositories, and what was found.

The instrument for this class is committed and no model has been called against it, so there is no table here. The fixtures, the answer key and the task set are in the repository and the projected size of the run is recorded beside them; an effect, an interval and a cost appear when the run happens and not before.

THE OWN SLOT FOR THIS CLASS IS EMPTY AND IS SAID TO BE EMPTY. THIS IS THE FIRST OWN SLOT IN THE COHORT THAT IS, and the reason is that the operator has no skill for this class. Every own slot before this one existed because the class was built around a skill the operator had already published; this class was built first and then looked for one. The repository at skills/ was read at the commit pinned on the code review slot, a67dd34c609f034c0cfd736a348659bbdf1605bf, and the two nearest candidates were weighed by the same standard the selection rule applies to an external slot, which is whether the when-to-use NAMES the class. skills/qa-testing names running QA on a page, a feature or a site at three depth tiers, after a deploy or before a launch, and hands code-level work to code-review-web in its own When NOT to use; it is verification of a running site, not the writing of a test for a module. skills/code-review-web names reviewing a pull request, debugging a production issue, investigating a build failure and auditing security, and hands pre-launch QA to qa-testing; it is review of code, not the writing of tests for it. No other skill in the repository names writing tests in any when-to-use: skills/frontend-component-build lists unit tests once as one row of a testing checklist for a component it is building, which is a step inside a different class rather than this one. Neither candidate is pinned, because a slot filled with the nearest thing to hand would make every number in this class a measurement of something other than what the class is named for.

Read for a skill naming this class, and none found: rampstackco/claude-skills. This class publishes NO row at all rather than one or a pair. It holds no claim, so nothing is planned, nothing is projected as a call against an arm, and no cell, effect, interval or verdict exists for it. What is committed is the instrument: the fixtures, the answer key, the runner and the task set. If the operator publishes a skill for this class, this slot is filled in its own change on the same footing as any other slot, and the external slot below is selected then.

THE EXTERNAL SLOT FOR THIS CLASS IS EMPTY AND IS SAID TO BE EMPTY. The selection rule is unchanged and is the one recorded with the cohort. It has not been run against these repositories for this class, so no candidate has been weighed, none is pinned and none is named here: naming a likely candidate before the rule had been applied to it would be a guess dressed as a record. The rule is run in its own change against a source list verified at the commit it names, and the candidates, their quoted sentences, their commits and the decision are recorded there on the same footing as every other external slot.

Repositories in scope, from the same selection rule as every other external slot: obra/superpowers, affaan-m/ECC, anthropics/skills. With the own slot for this class also empty, the class has no arm on either side. That is a different thing from a class with one arm and it is recorded as two entries rather than one, because the two are empty for different reasons and only one of them has had the rule applied to it.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.

Selection for the external slots follows one rule, recorded at harness/claims/skill-cohort.json and applied on 2026-08-26: From obra/superpowers, affaan-m/ECC and anthropics/skills, choose the skill whose when-to-use or equivalent most directly names the task class. If two qualify, take the one with the closer output format. If none qualifies in a repo, the slot may go to a second candidate from another listed repo. If no listed repo has a match for a class, the slot is left empty and said to be empty. A skill appears here as an identity pin and a measurement. Nothing else about it is published.

1 own slot is declared empty, for classes this site’s operator publishes no skill for. 1 external slot has been declared and not yet filled. A slot the cohort has declared and not yet filled. It is held here rather than in `slots` because a slot in `slots` carries a pin, and a pin is a commit: there is no honest pin for a selection that has not been made. Every gate that reads `slots` therefore sees the truth, and the class-and-pair invariant is stated over both lists together rather than relaxed. AMENDED 2026-08-31 BY THE SOURCE EXPANSION: THE CODE-REVIEW ENTRY IS GONE, BECAUSE THAT SELECTION HAS BEEN MADE. Seven candidates named the class, the output-format tiebreak was re-run against the basis committed with the class in PR #38, and under R23 all seven are pinned. The declaration it carried would now be false. WHAT REMAINS IS THE TEST-WRITING CLASS, on both sides, and it remains for a different reason from the one code-review had: not that a selection is pending, but that the operator publishes no skill naming the class and no listed repository holds one either. THE TWO KINDS OF EMPTINESS ARE NOT THE SAME and this list now holds only the second. Under R23 an entry here no longer means the one external slot is empty, because there is no longer a fixed number of external slots; it means the rule was run over the repositories in `searchScope` and produced nothing that both qualified and could be pinned. Everything the expansion could not pin is recorded where it belongs rather than nowhere: candidates barred by licence in `sourceRepos.barred` and `proposedSlotsNotTaken`, candidates that did not qualify as `qualifies: false` in each selection record with their bytes under `referencePins.rejected`, and re-hosts in `aggregators`.