Claude Skills Compared, Task by Task
Each skill here was tested on a fixed task set, once with the skill loaded into the model's context and once without. The tables show what changed. Nothing on this page describes what a skill is or says about itself; the commit link goes to the source.
How to read this
- Each skill was run on the same tasks with the skill loaded and without it, and where the third arm was run, with a one-line instruction naming the task instead.
- The sentence says what happened, against which baseline, and on which models.
- The bar shows what the model managed without the skill, and what the skill added.
- Loading a skill costs context on every request, whether or not it helps.
How to read the numbers
How to read these tables
- Without the skill
- The score on the task set before any skill was loaded. Where this is already near the top of the scale, there was little room for a skill to show an effect, and the result says so instead of reporting no difference.
- Without -> with
- The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
- Effect (range)
- With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
- Orbit
- Stable: delta at or above the pass threshold, and the interval excludes zero
- Past the horizon: delta at or below the failure floor, and the interval excludes zero
- In free drift: the interval spans zero, or the delta sits between the floor and the threshold
- Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound
A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.
- Cohort
- The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
- Commit
- The exact version tested. Later commits are not covered by this verdict.
- Injected context tokens
- How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
- Tested
- The date. Verdicts age as models change.
Skill authoring
The model is given a short brief and asked to write a complete SKILL.md. The score is the share of the open Agent Skills specification's requirements the result meets, counting only the 14 defined requirements that apply to the document written; requirements for optional fields the document does not use are not counted. A second reading against this site's own house style is shown on each skill's page but is not ranked, because one of the skills tested is ours.
5 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.
Every skill tested for writing skills, ranked.
Skill cohort (launch cells)
one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.
- Helped on someBest result in this test
skills/skill-creation-walkthrough rampstackco/claude-skills*
Helped clearly on Claude Haiku 4.5 and GPT-5 mini; helped a little on Gemini 3.1 Flash Lite, under our bar. Measured against no skill loaded.
Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · helped
- Gemini 3.1 Flash Lite · unclear. The model scored 87% without it.
- GPT-5 mini · helped
Adds about 8,400 tokens of context to every request.
- Helped on some
skills/skill-creator anthropics/skills
Helped clearly on Claude Haiku 4.5 and GPT-5 mini; helped a little on Gemini 3.1 Flash Lite, under our bar. Measured against no skill loaded. Based on 56 of 60 runs; the rest returned nothing.
Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · helped
- Gemini 3.1 Flash Lite · unclear. The model scored 87% without it.
- GPT-5 mini · helped
Adds about 11,500 tokens of context to every request.
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.
- Helped clearlyBest result in this test
skills/skill-creation-walkthrough rampstackco/claude-skills*
Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction. Based on 60 of 63 runs; the rest returned nothing.
Found 100% of the specification requirements, against 76% with no skill loaded. Against 68% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · helped
Adds about 8,200 tokens of context to every request.
- Helped clearly
skills/skill-creator anthropics/skills
Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction. Based on 60 of 112 runs; the rest returned nothing.
Found 100% of the specification requirements, against 82% with no skill loaded. Against 70% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · helped
Adds about 10,900 tokens of context to every request.
- Helped clearly
distribution/claude-plugin/skills/skill-builder yusufkaraaslan/Skill_Seekers
Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction. Based on 60 of 67 runs; the rest returned nothing.
Found 99% of the specification requirements, against 70% with no skill loaded. Against 73% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · helped
Adds about 1,200 tokens of context to every request.
- Helped on some
skills/skill-creator openclaw/openclaw
Helped on Gemini 3.1 Flash Lite; helped a little on GPT-5 mini, under our bar. Measured against the same task with a one-line instruction. Based on 60 of 74 runs; the rest returned nothing.
Found 99% of the specification requirements, against 75% with no skill loaded. Against 77% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · unclear. The model scored 63% without it.
Adds about 600 tokens of context to every request.
Show the numbers
| Skill | Cohort | Without the skill | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|---|
| No skill loaded Wider run, three ways | baseline | 76% | 76% | 0 by definition | 0 | 2026-09-05 | |
| One-line instruction instead Wider run, three ways | baseline | 72% | 0 by definition | 0 | 2026-09-05 | ||
skills/skill-creation-walkthrough from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this test | rampstackco/claude-skills* | 76% | 76% -> 100% | +0.32 [+0.25, +0.38] | Stable | 8,190 | 2026-09-05 |
skills/skill-creator from anthropics/skills 3b3fad96af16 Wider run, three ways | anthropics/skills | 82% | 82% -> 100% | +0.30 [+0.25, +0.34] | Stable | 10,872 | 2026-09-05 |
distribution/claude-plugin/skills/skill-builder from yusufkaraaslan/Skill_Seekers f3972efa33fa Wider run, three ways | yusufkaraaslan/Skill_Seekers | 70% | 70% -> 99% | +0.26 [+0.25, +0.27] | Stable | 1,203 | 2026-09-05 |
skills/skill-creator from openclaw/openclaw 61b08f0ebb7a Wider run, three ways | openclaw/openclaw | 75% | 75% -> 99% | +0.22 [+0.19, +0.25] | Mixed, see page | 587 | 2026-09-05 |
| No skill loaded First run, with and without the skill | baseline | 74% | 74% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead First run, with and without the skill | baseline | instruction arm: not measured | 0 by definition | 0 | 2026-08-30 | ||
skills/skill-creation-walkthrough from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this test | rampstackco/claude-skills* | 74% | 74% -> 100% | +0.26 [+0.13, +0.38] | Stable on 2 of 3 | 8,390 | 2026-08-30 |
skills/skill-creator from anthropics/skills 3b3fad96af16 First run, with and without the skill | anthropics/skills | 74% | 74% -> 100% | +0.26 [+0.13, +0.38] | Stable on 2 of 3 | 11,514 | 2026-08-30 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | |||||||
In Claude Code
These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.
Skill cohort (one pass)
one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.
skills/skill-creation-walkthrough
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · could not be measured, and both runs agree
skills/skill-creator
- Claude Fable 5 · could not be measured, and both runs agree
- Claude Opus 5 · unclear under repeat sampling; could not be measured in the single-pass run
Panel v2 (repeat sampling)
two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.
skills/skill-creation-walkthrough
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · could not be measured, and both runs agree
- Claude Haiku 4.5 · helped under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
skills/skill-creator
- Claude Fable 5 · could not be measured, and both runs agree
- Claude Opus 5 · unclear under repeat sampling; could not be measured in the single-pass run
- Claude Haiku 4.5 · helped under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
Fable 5.1, one pass
one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.
skills/skill-creation-walkthrough
- Claude Fable 5.1 · could not be measured in the version re-test
skills/skill-creator
- Claude Fable 5.1 · made results worse in the version re-test
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
skills/skill-creation-walkthrough
- Claude Opus 5 · could not be measured, and both runs agree
- Claude Haiku 4.5 · helped under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · could not be measured in the version re-test
skills/skill-creator
- Claude Opus 5 · unclear under repeat sampling; could not be measured in the single-pass run
- Claude Haiku 4.5 · helped under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · made results worse in the version re-test
skills/skill-creator
- Claude Fable 5.1 · could not be measured in the three-arm cohort
- Claude Haiku 4.5 · helped in the three-arm cohort
- Claude Opus 5 · could not be measured in the three-arm cohort
- Claude Sonnet 5 · could not be measured in the three-arm cohort
distribution/claude-plugin/skills/skill-builder
- Claude Fable 5.1 · could not be measured in the three-arm cohort
- Claude Haiku 4.5 · helped in the three-arm cohort
- Claude Opus 5 · could not be measured in the three-arm cohort
- Claude Sonnet 5 · could not be measured in the three-arm cohort
| Skill | Claude Haiku 4.5 · skill-cohort · | Gemini 3.1 Flash Lite · skill-cohort · | GPT-5 mini · skill-cohort · | Gemini 3.1 Flash Lite · expansion-cohort · | GPT-5 mini · expansion-cohort · |
|---|---|---|---|---|---|
| rampstackco-skill-creation-walkthrough · Wider run, three ways | Not run | Not run | Not run | helped · +0.255 | helped · +0.382 |
| anthropics-skill-creator · Wider run, three ways | Not run | Not run | Not run | helped · +0.255 | helped · +0.336 |
| rampstackco-skill-creation-walkthrough · First run, with and without the skill | helped · +0.382 | unclear · +0.127 | helped · +0.282 | Not run | Not run |
| yusufkaraaslan-skill-builder · Wider run, three ways | Not run | Not run | Not run | helped · +0.255 | helped · +0.273 |
| anthropics-skill-creator · First run, with and without the skill | helped · +0.382 | unclear · +0.127 | helped · +0.273 | Not run | Not run |
| openclaw-skill-creator · Wider run, three ways | Not run | Not run | Not run | helped · +0.255 | unclear · +0.192 |
Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.
How this external slot was filled
Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in anthropics/skills skills/skill-creator at 3b3fad96af16.
Two candidates name the class. The task class is scored on a single SKILL.md whose frontmatter must parse and whose section order must match SKILL_AUTHORING.md. writing-skills prescribes that document's section order under the heading above. skill-creator prescribes an authoring and evaluation workflow, and the structure it prescribes under its own Report structure heading is an evaluation report rather than the SKILL.md. writing-skills is the closer output format.
Addendum. RECORDED AFTER THE FACT, ON 2026-08-26, AND NOT ACTED ON. This tiebreak was decided when the skill-authoring class was scored against SKILL_AUTHORING.md alone. That document is now Key B, disclosed and unranked, and the ranked key is the open Agent Skills specification. The tiebreak reasoning above therefore turns on a basis that is no longer the ranked one. It is left exactly as it was decided, because a selection record that is rewritten to agree with a later ruling is not a record of what was decided. The selection was NOT re-run against Key A, and whether it would produce the same candidate is an open question stated here rather than assumed.
Second addendum. SECOND ADDENDUM, 2026-08-26. THE TIEBREAK WAS RE-RUN WITH KEY A AS THE OUTPUT-FORMAT BASIS, AND IT SWAPPED. Basis: the selection rule is unchanged, but the output format the class is scored on is now the open Agent Skills specification rather than SKILL_AUTHORING.md. Candidates compared: the same two that qualified, obra/superpowers skills/writing-skills and anthropics/skills skills/skill-creator. Result: skill-creator. It prescribes the specification's own model of the artifact, naming the required and the optional frontmatter fields, the bundled directory layout, the progressive disclosure levels and the line ceiling. writing-skills prescribes a body section order, which is the one part of a SKILL.md the specification explicitly leaves unconstrained, and it disagrees with the specification in three places: it puts the thousand and twenty four character limit on the whole frontmatter rather than on the description, it tells an author the description must not say what the skill does where the specification asks for both, and the example name in its own template is capitalised where the specification allows lowercase only. Run through Key A, writing-skills' own template fails a predicate and skill-creator's passes every applicable one. The original record and the first addendum above are left exactly as written. obra/superpowers skills/writing-skills moves to the rejected pins, where its bytes stay committed and its quoted sentence stays resolvable.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/writing-skills from obra/superpowers | b36e0829c6d0 | d34db5c8aed6 | yes | Names the class in its when-to-use directly. |
| skills/skill-creator from anthropics/skills | 3b3fad96af16 | e51ce07ea7fd | yes | Also names the class directly, so the tiebreak applies. |
| skills/skill-scout from affaan-m/ECC | d8409a4b0813 | 71b1e8a4aede | no | Searches for an existing skill before one is written. It names the moment before authoring, not authoring. |
| skills/skill-comply from affaan-m/ECC | d8409a4b0813 | e26cbbc30a32 | no | Measures whether a written skill is obeyed. It is about a skill, not about writing one. |
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in openclaw/openclaw skills/skill-creator at 61b08f0ebb7a.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/skill-creator from openclaw/openclaw | 61b08f0ebb7a | eea1a8321d20 | yes | Names authoring and reviewing SKILL.md files as such, with no domain narrowing. Pinned. |
| distribution/claude-plugin/skills/skill-builder from yusufkaraaslan/Skill_Seekers | f3972efa33fa | eea1a8321d20 | yes | Names creating skills in its when-to-use. Its inputs are source materials rather than an authoring brief, which is where it sits relative to the others; under R23 that is ordering and not exclusion. Pinned. |
| skills/microsoft-skill-creator from github/awesome-copilot | c956566a35c3 | 2b1e9f91f708 | no | Names the class and then narrows it to one vendor's technologies, and declares a network dependency on the Learn MCP server. Rejected on the same ground as affaan-m/ECC flutter-dart-code-review and wshobson/agents screen-reader-testing: a class-general skill bound to one vendor's surface is not the class. NOT a tiebreak loss. |
- addyosmani/agent-skills skills/using-agent-skills names discovering and invoking skills, not authoring one, the same ground affaan-m/ECC skills/skill-scout was rejected on in the committed record.
- browser-act/skills browser-act-skill-forge names forging a SKILL.md and conditions it throughout on website exploration and bulk extraction.
- mattpocock/skills has no skill whose when-to-use names authoring a skill.
- wshobson/agents has none: its skill-forge-essentials plugin holds ai-debt-detector, session-guard and visual-edit-precision, and none names authoring.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- alirezarezvani/claude-skills has none.
- Jeffallan/claude-skills has none.
- EveryInc/compound-engineering-plugin has none.
- open-gsd/gsd-core has none.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in yusufkaraaslan/Skill_Seekers distribution/claude-plugin/skills/skill-builder at f3972efa33fa.
PINNED UNDER R23, HAVING QUALIFIED AND BEEN DISPLACED BY THE CAP. This skill appears in the recon's skill-authoring qualifying table and was not proposed for a slot, because the class could hold one external skill and openclaw/openclaw skill-creator was the closer match: it names authoring and reviewing SKILL.md files as such, where this one's inputs are documentation, repositories, PDFs and videos, so its subject is building a skill FROM source material rather than authoring one to a brief. That is still the honest comparison and it is no longer a reason for absence.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/skill-creator from openclaw/openclaw | 61b08f0ebb7a | 42c7e6285ae8 | yes | Names authoring and reviewing SKILL.md files as such, with no domain narrowing. Pinned. |
| distribution/claude-plugin/skills/skill-builder from yusufkaraaslan/Skill_Seekers | f3972efa33fa | 42c7e6285ae8 | yes | Names creating skills in its when-to-use. Its inputs are source materials rather than an authoring brief, which is where it sits relative to the others; under R23 that is ordering and not exclusion. Pinned. |
| skills/microsoft-skill-creator from github/awesome-copilot | c956566a35c3 | 2b1e9f91f708 | no | Names the class and then narrows it to one vendor's technologies, and declares a network dependency on the Learn MCP server. Rejected on the same ground as affaan-m/ECC flutter-dart-code-review and wshobson/agents screen-reader-testing: a class-general skill bound to one vendor's surface is not the class. NOT a tiebreak loss. |
- addyosmani/agent-skills skills/using-agent-skills names discovering and invoking skills, not authoring one, the same ground affaan-m/ECC skills/skill-scout was rejected on in the committed record.
- browser-act/skills browser-act-skill-forge names forging a SKILL.md and conditions it throughout on website exploration and bulk extraction.
- mattpocock/skills has no skill whose when-to-use names authoring a skill.
- wshobson/agents has none: its skill-forge-essentials plugin holds ai-debt-detector, session-guard and visual-edit-precision, and none names authoring.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- alirezarezvani/claude-skills has none.
- Jeffallan/claude-skills has none.
- EveryInc/compound-engineering-plugin has none.
- open-gsd/gsd-core has none.
Accessibility audit
10 fixture web pages, each seeded with known WCAG defects, 48 in total. The model is asked to list the defects it finds; the score is the share of seeded defects it names. 23 defects are verified in the fixtures by mechanical check; 25 are judgement items the check cannot verify and are listed as such.
4 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.
Every skill tested for accessibility, ranked.
Skill cohort (launch cells)
one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.
- Helped on someBest result in this test
skills/accessibility-audit rampstackco/claude-skills*
Helped on Claude Haiku 4.5; no reliable difference on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.
Found 71% of the seeded accessibility defects, against 58% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · helped
- Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
- GPT-5 mini · unclear. The model scored 69% without it.
Adds about 11,700 tokens of context to every request.
- No reliable difference
skills/accessibility affaan-m/ECC
No reliable difference on any model tested. The models scored 58% without it. Measured against no skill loaded.
Found 56% of the seeded accessibility defects, against 58% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · unclear. The model scored 39% without it.
- Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
- GPT-5 mini · unclear. The model scored 71% without it.
Adds about 2,200 tokens of context to every request.
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.
- No reliable differenceBest result in this test
skills/accessibility-audit rampstackco/claude-skills*
No reliable difference on any model tested. The models scored 60% without it. Measured against the same task with a one-line instruction.
Found 68% of the seeded accessibility defects, against 60% with no skill loaded. Against 60% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
- GPT-5 mini · unclear. The model scored 57% without it.
Adds about 11,300 tokens of context to every request.
- No reliable difference
plugins/accessibility-compliance/skills/wcag-audit-patterns wshobson/agents
No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.
Found 63% of the seeded accessibility defects, against 63% with no skill loaded. Against 58% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
- GPT-5 mini · unclear. The model scored 64% without it.
Adds about 4,000 tokens of context to every request.
- No reliable difference
engineering-team/a11y-audit/skills/a11y-audit alirezarezvani/claude-skills
No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction. Based on 60 of 62 runs; the rest returned nothing.
Found 65% of the seeded accessibility defects, against 64% with no skill loaded. Against 62% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
- GPT-5 mini · unclear. The model scored 64% without it.
Adds about 16,400 tokens of context to every request.
- No reliable difference
skills/accessibility affaan-m/ECC
No reliable difference on any model tested. The models scored 55% without it. Measured against the same task with a one-line instruction. Based on 60 of 61 runs; the rest returned nothing.
Found 57% of the seeded accessibility defects, against 55% with no skill loaded. Against 64% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 63% without it.
- GPT-5 mini · unclear. The model scored 48% without it.
Adds about 2,100 tokens of context to every request.
Show the numbers
| Skill | Cohort | Without the skill | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|---|
| No skill loaded First run, with and without the skill | baseline | 58% | 58% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead First run, with and without the skill | baseline | instruction arm: not measured | 0 by definition | 0 | 2026-08-30 | ||
skills/accessibility-audit from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this test | rampstackco/claude-skills* | 58% | 58% -> 71% | +0.13 [-0.00, +0.29] | In free drift on 2 of 3 | 11,729 | 2026-08-30 |
skills/accessibility from affaan-m/ECC d8409a4b0813 First run, with and without the skill | affaan-m/ECC | 58% | 58% -> 56% | -0.02 [-0.10, +0.13] | In free drift | 2,206 | 2026-08-30 |
| No skill loaded Wider run, three ways | baseline | 61% | 61% | 0 by definition | 0 | 2026-09-05 | |
| One-line instruction instead Wider run, three ways | baseline | 61% | 0 by definition | 0 | 2026-09-05 | ||
skills/accessibility-audit from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this test | rampstackco/claude-skills* | 60% | 60% -> 68% | +0.08 [+0.05, +0.10] | In free drift | 11,295 | 2026-09-05 |
plugins/accessibility-compliance/skills/wcag-audit-patterns from wshobson/agents 38e19c20d2b1 Wider run, three ways | wshobson/agents | 63% | 63% -> 63% | +0.05 [+0.02, +0.09] | In free drift | 3,967 | 2026-09-05 |
engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills 19392f7a0826 Wider run, three ways | alirezarezvani/claude-skills | 64% | 64% -> 65% | +0.03 [-0.03, +0.09] | In free drift | 16,407 | 2026-09-05 |
skills/accessibility from affaan-m/ECC d8409a4b0813 Wider run, three ways | affaan-m/ECC | 55% | 55% -> 57% | -0.06 [-0.13, +0.01] | In free drift | 2,108 | 2026-09-05 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | |||||||
In Claude Code
These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.
Skill cohort (one pass)
one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.
skills/accessibility-audit
- Claude Opus 5 · unclear, and both runs agree
skills/accessibility
- Claude Opus 5 · unclear, and both runs agree
Verdict withheld on this run
skills/accessibility-audit, Claude Fable 5: served model mismatch; 1 usable pair. Kept and not scored into any cell.
skills/accessibility, Claude Fable 5: served model mismatch; 2 usable pairs. Kept and not scored into any cell.
Panel v2 (repeat sampling)
two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.
skills/accessibility-audit
- Claude Opus 5 · unclear, and both runs agree
- Claude Fable 5 · unclear under repeat sampling
- Claude Haiku 4.5 · helped under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
skills/accessibility
- Claude Opus 5 · unclear, and both runs agree
- Claude Fable 5 · unclear under repeat sampling
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
Fable 5.1, one pass
one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.
skills/accessibility-audit
- Claude Fable 5.1 · unclear in the version re-test
skills/accessibility
- Claude Fable 5.1 · unclear in the version re-test
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
skills/accessibility-audit
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · helped under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · unclear in the version re-test
skills/accessibility
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · unclear in the version re-test
plugins/accessibility-compliance/skills/wcag-audit-patterns
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
engineering-team/a11y-audit/skills/a11y-audit
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · helped in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
| Skill | Claude Haiku 4.5 · skill-cohort · | Gemini 3.1 Flash Lite · skill-cohort · | GPT-5 mini · skill-cohort · | Gemini 3.1 Flash Lite · expansion-cohort · | GPT-5 mini · expansion-cohort · |
|---|---|---|---|---|---|
| rampstackco-accessibility-audit · First run, with and without the skill | helped · +0.289 | unclear · +0.094 | unclear · -0.003 | Not run | Not run |
| rampstackco-accessibility-audit · Wider run, three ways | Not run | Not run | Not run | unclear · +0.054 | unclear · +0.097 |
| wshobson-wcag-audit-patterns · Wider run, three ways | Not run | Not run | Not run | unclear · +0.017 | unclear · +0.087 |
| alirezarezvani-a11y-audit · Wider run, three ways | Not run | Not run | Not run | unclear · +0.090 | unclear · -0.026 |
| affaan-m-accessibility · First run, with and without the skill | unclear · +0.131 | unclear · -0.095 | unclear · -0.083 | Not run | Not run |
| affaan-m-accessibility · Wider run, three ways | Not run | Not run | Not run | unclear · -0.135 | unclear · +0.012 |
Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/accessibility at d8409a4b0813.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/accessibility from affaan-m/ECC | d8409a4b0813 | ab86a50717d6 | yes | Names auditing against WCAG in the when-to-use itself. |
| skills/click-path-audit from affaan-m/ECC | d8409a4b0813 | 662fc5d02483 | no | An audit of behavioural state, not of accessibility. The word audit is shared and the class is not. |
- obra/superpowers has no skill naming accessibility in any when-to-use.
- anthropics/skills has none either; webapp-testing names browser testing rather than accessibility.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in wshobson/agents plugins/accessibility-compliance/skills/wcag-audit-patterns at 38e19c20d2b1.
TWO SKILLS IN ONE REPOSITORY, AND THE COMPARISON BETWEEN THEM IS NOT A TIEBREAK. wshobson/agents also carries plugins/ui-design/skills/accessibility-compliance, which the recon's table lists as qualifying. It is recorded here as NOT qualifying, and that is a judgment about the class rather than a ranking: its deliverable is a compliant interface, it reaches auditing only as the first of four alternatives, and the artifact this class scores is a findings record of locators against success criteria, which a skill that produces code does not produce. R23 removed the per-class cap, so nothing about that decision turns on there being room for one skill: were it a tiebreak loss it would be pinned alongside this slot like every other qualifying candidate. It is a rejection, its bytes are committed under referencePins.rejected, and its sentence is quoted there.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| plugins/accessibility-compliance/skills/wcag-audit-patterns from wshobson/agents | 38e19c20d2b1 | 7c2619d811fe | yes | The first clause is the class label. The strongest names-the-class match found anywhere in the pass, for any class. Pinned. |
| engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills | 19392f7a0826 | 7c2619d811fe | yes | Names auditing accessibility in the first clause of its when-to-use. Pinned. |
| plugins/ui-design/skills/accessibility-compliance from wshobson/agents | 38e19c20d2b1 | 3ad0fbcb0a69 | no | DOES NOT QUALIFY, AND THIS IS A QUALIFICATION JUDGMENT RATHER THAN A TIEBREAK LOSS. Its deliverable is a compliant interface: it leads with implementing interfaces with inclusive design and assistive-technology support, and reaches auditing only as the first of four alternatives. The class is an audit, and the artifact this class scores is a findings record of locators against success criteria, which a skill that produces code does not produce. R23 removed the cap and did not turn this into an admission: it was never displaced by a better candidate, it was judged not to name the class. |
| plugins/accessibility-compliance/skills/screen-reader-testing from wshobson/agents | 38e19c20d2b1 | 591c7ff80f40 | no | One assistive-technology modality tested directly. Not an audit against a standard. |
- mattpocock/skills has no skill whose when-to-use names accessibility.
- addyosmani/agent-skills has none.
- openclaw/openclaw has none.
- github/awesome-copilot has none.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- Jeffallan/claude-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in alirezarezvani/claude-skills engineering-team/a11y-audit/skills/a11y-audit at 19392f7a0826.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| plugins/accessibility-compliance/skills/wcag-audit-patterns from wshobson/agents | 38e19c20d2b1 | 34f619aeb9cc | yes | The first clause is the class label. The strongest names-the-class match found anywhere in the pass, for any class. Pinned. |
| engineering-team/a11y-audit/skills/a11y-audit from alirezarezvani/claude-skills | 19392f7a0826 | 34f619aeb9cc | yes | Names auditing accessibility in the first clause of its when-to-use. Pinned. |
| plugins/ui-design/skills/accessibility-compliance from wshobson/agents | 38e19c20d2b1 | 3ad0fbcb0a69 | no | DOES NOT QUALIFY, AND THIS IS A QUALIFICATION JUDGMENT RATHER THAN A TIEBREAK LOSS. Its deliverable is a compliant interface: it leads with implementing interfaces with inclusive design and assistive-technology support, and reaches auditing only as the first of four alternatives. The class is an audit, and the artifact this class scores is a findings record of locators against success criteria, which a skill that produces code does not produce. R23 removed the cap and did not turn this into an admission: it was never displaced by a better candidate, it was judged not to name the class. |
| plugins/accessibility-compliance/skills/screen-reader-testing from wshobson/agents | 38e19c20d2b1 | 591c7ff80f40 | no | One assistive-technology modality tested directly. Not an audit against a standard. |
- mattpocock/skills has no skill whose when-to-use names accessibility.
- addyosmani/agent-skills has none.
- openclaw/openclaw has none.
- github/awesome-copilot has none.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- Jeffallan/claude-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- yusufkaraaslan/Skill_Seekers has none.
Spec writing
The model is given a feature request and asked for a product spec meeting a stated contract: required sections and testable acceptance criteria. Scored two ways: against the contract, and by a grader comparing the with and without outputs blind. Both readings appear.
7 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.
Every skill tested for product management, ranked.
Skill cohort (launch cells)
one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.
- Helped on someBest result in this test
skills/product-capability affaan-m/ECC
Helped clearly on Claude Haiku 4.5 and Gemini 3.1 Flash Lite; unclear on GPT-5 mini. Measured against no skill loaded.
A judge preferred its output 27 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.
- Claude Haiku 4.5 · helped
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · unclear. The model scored 30% without it.
Adds about 1,100 tokens of context to every request.
- Helped clearly
skills/pm-spec-writing rampstackco/claude-skills*
Helped clearly on Claude Haiku 4.5, Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.
A judge preferred its output 25 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.
- Claude Haiku 4.5 · helped
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · helped
Adds about 6,800 tokens of context to every request.
- Nothing left to measureBest result in this test
skills/product-capability affaan-m/ECC
Nothing left to measure. The models scored 98% without the skill. Measured against no skill loaded. Models already did well without a skill, so there was little room to improve.
Found 95% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · Nothing left to measure. The model scored 100% without the skill.
- Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
- GPT-5 mini · Nothing left to measure. The model scored 94% without the skill.
Adds about 1,100 tokens of context to every request.
- Nothing left to measure
skills/pm-spec-writing rampstackco/claude-skills*
Nothing left to measure. The models scored 98% without the skill. Measured against no skill loaded. Models already did well without a skill, so there was little room to improve.
Found 94% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · Nothing left to measure. The model scored 99% without the skill.
- Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
- GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.
Adds about 6,800 tokens of context to every request.
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.
- Nothing left to measureBest result in this test
skills/product-capability affaan-m/ECC
Nothing left to measure. The models scored 98% without the skill. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.
Found 98% of the required sections and acceptance criteria, against 98% with no skill loaded. Against 97% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
- GPT-5 mini · Nothing left to measure. The model scored 97% without the skill.
Adds about 1,000 tokens of context to every request.
- Could not be measured
skills/spec-driven-development addyosmani/agent-skills
Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.
Found 98% of the required sections and acceptance criteria, against 98% with no skill loaded. Against 98% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · not measured
- GPT-5 mini · Nothing left to measure. The model scored 96% without the skill.
Adds about 2,800 tokens of context to every request.
- Could not be measured
pm-execution/skills/create-prd phuryn/pm-skills
Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.
Found 98% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 98% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · not measured
- GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.
Adds about 900 tokens of context to every request.
- Could not be measured
engineering/skills/spec-driven-workflow alirezarezvani/claude-skills
Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.
Found 97% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 97% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · not measured
- GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.
Adds about 15,000 tokens of context to every request.
- Could not be measured
skills/prd github/awesome-copilot
Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve. Based on 60 of 62 runs; the rest returned nothing.
Found 95% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 98% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · not measured
- GPT-5 mini · Nothing left to measure. The model scored 95% without the skill.
Adds about 1,100 tokens of context to every request.
- Nothing left to measure
skills/pm-spec-writing rampstackco/claude-skills*
Nothing left to measure. The models scored 99% without the skill. Measured against the same task with a one-line instruction. Models already did well without a skill, so there was little room to improve.
Found 95% of the required sections and acceptance criteria, against 99% with no skill loaded. Against 98% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · Nothing left to measure. The model scored 99% without the skill.
- GPT-5 mini · Nothing left to measure. The model scored 99% without the skill.
Adds about 6,700 tokens of context to every request.
Show the numbers
| Skill | Cohort | Without the skill | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|---|
| No skill loaded Blind comparison of the same task, run twice, with ties counting half, First run, with and without the skill | baseline | 13% | 13% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead Blind comparison of the same task, run twice, with ties counting half, First run, with and without the skill | baseline | instruction arm: not applicable to a blind preference | 0 by definition | 0 | 2026-08-30 | ||
skills/product-capability from affaan-m/ECC d8409a4b0813 Blind comparison of the same task, run twice, with ties counting half First run, with and without the skill Best result in this test | affaan-m/ECC | 10% | 10% -> 90% baseline shown as the complement | +0.80 [+0.40, +1.00] | Stable on 2 of 3 | 1,082 | 2026-08-30 |
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Blind comparison of the same task, run twice, with ties counting half First run, with and without the skill | rampstackco/claude-skills* | 17% | 17% -> 83% baseline shown as the complement | +0.67 [+0.60, +0.80] | Stable | 6,845 | 2026-08-30 |
| No skill loaded Deterministic, scored against a committed answer key, Wider run, three ways | baseline | 98% | 98% | 0 by definition | 0 | 2026-09-05 | |
| One-line instruction instead Deterministic, scored against a committed answer key, Wider run, three ways | baseline | 98% | 0 by definition | 0 | 2026-09-05 | ||
skills/product-capability from affaan-m/ECC d8409a4b0813 Deterministic, scored against a committed answer key Wider run, three ways Best result in this test | affaan-m/ECC | 98% at the ceiling | 98% -> 98% | +0.01 [-0.01, +0.03] | In free drift | 1,047 | 2026-09-05 |
skills/spec-driven-development from addyosmani/agent-skills d2c37ef6225d Deterministic, scored against a committed answer key Wider run, three ways | addyosmani/agent-skills | 98% at the ceiling | 98% -> 98% | +0.00 [+0.00, +0.00] | Mixed, see page | 2,777 | 2026-09-05 |
pm-execution/skills/create-prd from phuryn/pm-skills 18468a95b427 Deterministic, scored against a committed answer key Wider run, three ways | phuryn/pm-skills | 97% at the ceiling | 97% -> 98% | +0.00 [+0.00, +0.00] | Mixed, see page | 941 | 2026-09-05 |
engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills 19392f7a0826 Deterministic, scored against a committed answer key Wider run, three ways | alirezarezvani/claude-skills | 97% at the ceiling | 97% -> 97% | -0.00 [-0.01, +0.00] | Mixed, see page | 15,043 | 2026-09-05 |
skills/prd from github/awesome-copilot c956566a35c3 Deterministic, scored against a committed answer key Wider run, three ways | github/awesome-copilot | 97% at the ceiling | 97% -> 95% | -0.03 [-0.06, +0.00] | Mixed, see page | 1,147 | 2026-09-05 |
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Deterministic, scored against a committed answer key Wider run, three ways | rampstackco/claude-skills* | 99% at the ceiling | 99% -> 95% | -0.04 [-0.06, -0.01] | In free drift | 6,653 | 2026-09-05 |
| No skill loaded Deterministic, scored against a committed answer key, First run, with and without the skill | baseline | 98% | 98% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead Deterministic, scored against a committed answer key, First run, with and without the skill | baseline | instruction arm: not measured | 0 by definition | 0 | 2026-08-30 | ||
skills/product-capability from affaan-m/ECC d8409a4b0813 Deterministic, scored against a committed answer key First run, with and without the skill Best result in this test | affaan-m/ECC | 98% at the ceiling | 98% -> 95% | -0.02 [-0.04, +0.00] | In free drift | 1,082 | 2026-08-30 |
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Deterministic, scored against a committed answer key First run, with and without the skill | rampstackco/claude-skills* | 98% at the ceiling | 98% -> 94% | -0.04 [-0.12, +0.00] | In free drift | 6,845 | 2026-08-30 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | |||||||
In Claude Code
These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.
Skill cohort (one pass)
one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.
skills/pm-spec-writing
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
skills/product-capability
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
Panel v2 (repeat sampling)
two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.
skills/pm-spec-writing
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
skills/product-capability
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
Fable 5.1, one pass
one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.
skills/pm-spec-writing
- Claude Fable 5.1 · unclear in the version re-test
skills/product-capability
- Claude Fable 5.1 · unclear in the version re-test
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
skills/pm-spec-writing
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · unclear in the version re-test
skills/product-capability
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · unclear in the version re-test
pm-execution/skills/create-prd
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
engineering/skills/spec-driven-workflow
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
skills/spec-driven-development
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
skills/prd
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.
How this external slot was filled
Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in affaan-m/ECC skills/product-capability at d8409a4b0813.
Two candidates name the class. The task class is scored on a document with required sections and acceptance criteria, so the closer output format is the one that fixes its sections. product-capability declares a canonical artifact and a section-by-section output format under the heading above. doc-coauthoring prescribes a three-stage collaboration and fixes no sections in the document it produces.
Addendum. RECORDED AFTER THE FACT, ON 2026-08-26, AND NOT ACTED ON. This tiebreak was decided when the skill-authoring class was scored against SKILL_AUTHORING.md alone. That document is now Key B, disclosed and unranked, and the ranked key is the open Agent Skills specification. The tiebreak reasoning above therefore turns on a basis that is no longer the ranked one. It is left exactly as it was decided, because a selection record that is rewritten to agree with a later ruling is not a record of what was decided. The selection was NOT re-run against Key A, and whether it would produce the same candidate is an open question stated here rather than assumed.
Second addendum. SECOND ADDENDUM, 2026-08-26. RE-EXAMINED, NOT RE-RUN, AND THE CANDIDATE IS UNCHANGED. Basis: this tiebreak was never decided on the operator's house standard. The spec-writing class is scored on the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms, which is a contract the two arms are given rather than a convention one of them was taught, and R7 did not move it. Key A is the answer key for the skill-authoring class and has no bearing here. Candidates compared: the same two that qualified, affaan-m/ECC skills/product-capability and anthropics/skills skills/doc-coauthoring. Result: unchanged, product-capability. THE FIRST ADDENDUM ABOVE IS WRONG ON THIS SLOT and is left in place rather than edited. It was applied to both output-format tiebreaks at once and says this one was decided against SKILL_AUTHORING.md, which it was not. Correcting it by rewriting would hide that the error was made; this sentence is the correction.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/product-capability from affaan-m/ECC | d8409a4b0813 | 3e0052802969 | yes | Names writing a specification from product intent. |
| skills/doc-coauthoring from anthropics/skills | 3b3fad96af16 | 2e47d78846fa | yes | Names technical specs among several document kinds, so the tiebreak applies. |
| skills/writing-plans from obra/superpowers | b36e0829c6d0 | 48508f44bbfd | no | Takes a spec as its input. It is downstream of the class, not the class. |
| skills/product-lens from affaan-m/ECC | d8409a4b0813 | d082be7c3dd9 | no | Rules itself out in its own words and hands the class to product-capability. |
How this external slot was filled
Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md frontmatter, description, in phuryn/pm-skills pm-execution/skills/create-prd at 18468a95b427.
THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| pm-execution/skills/create-prd from phuryn/pm-skills | 18468a95b427 | 6f9493aa031b | yes | Fixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned. |
| engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills | 19392f7a0826 | 6f9493aa031b | yes | The only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned. |
| skills/spec-driven-development from addyosmani/agent-skills | d2c37ef6225d | 6f9493aa031b | yes | Names the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits. |
| skills/prd from github/awesome-copilot | c956566a35c3 | 6f9493aa031b | yes | Lists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23. |
| skills/spec-miner from Jeffallan/claude-skills | 882ef55e377d | 0deb55afcc6c | no | Its input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class. |
- mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
- openclaw/openclaw has no skill whose when-to-use names writing a specification.
- samber/cc-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- wshobson/agents has none.
- yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled
Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md frontmatter, description, in alirezarezvani/claude-skills engineering/skills/spec-driven-workflow at 19392f7a0826.
THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| pm-execution/skills/create-prd from phuryn/pm-skills | 18468a95b427 | 119991870509 | yes | Fixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned. |
| engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills | 19392f7a0826 | 119991870509 | yes | The only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned. |
| skills/spec-driven-development from addyosmani/agent-skills | d2c37ef6225d | 119991870509 | yes | Names the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits. |
| skills/prd from github/awesome-copilot | c956566a35c3 | 119991870509 | yes | Lists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23. |
| skills/spec-miner from Jeffallan/claude-skills | 882ef55e377d | 0deb55afcc6c | no | Its input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class. |
- mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
- openclaw/openclaw has no skill whose when-to-use names writing a specification.
- samber/cc-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- wshobson/agents has none.
- yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in addyosmani/agent-skills skills/spec-driven-development at d2c37ef6225d.
THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| pm-execution/skills/create-prd from phuryn/pm-skills | 18468a95b427 | 280779b1914e | yes | Fixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned. |
| engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills | 19392f7a0826 | 280779b1914e | yes | The only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned. |
| skills/spec-driven-development from addyosmani/agent-skills | d2c37ef6225d | 280779b1914e | yes | Names the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits. |
| skills/prd from github/awesome-copilot | c956566a35c3 | 280779b1914e | yes | Lists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23. |
| skills/spec-miner from Jeffallan/claude-skills | 882ef55e377d | 0deb55afcc6c | no | Its input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class. |
- mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
- openclaw/openclaw has no skill whose when-to-use names writing a specification.
- samber/cc-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- wshobson/agents has none.
- yusufkaraaslan/Skill_Seekers has none.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in github/awesome-copilot skills/prd at c956566a35c3.
THE COMMITTED BASIS FOR THIS CLASS is the one already recorded at `tiebreakAddendum2` on affaan-m-product-capability: the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms. Against it, create-prd fixes a named 8-section template, spec-driven-workflow is the only candidate naming acceptance criteria in those words, spec-driven-development names the class and fixes neither half, and awesome-copilot prd lists contents without fixing an order. UNDER R23 THIS ANALYSIS ORDERS THE RECORD AND ADMITS NOBODY AND EXCLUDES NOBODY. It was written when a class could hold one external skill, so the comparison below decided which single candidate was pinned. The cap is gone: every candidate that qualified and whose bytes this project may hold is pinned, and this slot is one of them. The comparison is kept because it is what was weighed and on which sentence, which is what R2 requires, and because how the candidates sit relative to one another is a real finding about the field. It is no longer a reason anybody is absent.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| pm-execution/skills/create-prd from phuryn/pm-skills | 18468a95b427 | fc1caf7b8c27 | yes | Fixes a named 8-section template in its own description: the closest match to a required-section list found in the pass, which is one half of this class's committed basis. Pinned. |
| engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills | 19392f7a0826 | fc1caf7b8c27 | yes | The only candidate whose when-to-use names acceptance criteria in those words, which is the other half of the committed basis. Pinned. |
| skills/spec-driven-development from addyosmani/agent-skills | d2c37ef6225d | fc1caf7b8c27 | yes | Names the class cleanly and fixes neither half of the basis: no section list, no acceptance criteria. It decomposes into a capability map of modules. Under the cap it was third on output format and unpinned; under R23 it is pinned and third is a description of where it sits. |
| skills/prd from github/awesome-copilot | c956566a35c3 | fc1caf7b8c27 | yes | Lists its contents as an inventory of what is included rather than as a fixed section order, so it does not fix the section list the basis turns on. Pinned under R23. |
| skills/spec-miner from Jeffallan/claude-skills | 882ef55e377d | 0deb55afcc6c | no | Its input is code and its output documents what already exists. The class writes a spec from product intent, before the code. The mirror image of the class, not the class. |
- mattpocock/skills skills/engineering/to-spec names producing a spec but publishes it to an issue tracker, carries `disable-model-invocation: true`, and requires prior state, so its output is a tracker issue rather than a sectioned document and it is not self-contained enough to pin.
- openclaw/openclaw has no skill whose when-to-use names writing a specification.
- samber/cc-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- wshobson/agents has none.
- yusufkaraaslan/Skill_Seekers has none.
On-page audit
10 fixture pages seeded with 43 known on-page SEO issues from a closed list of 19 issue types. The model is asked to list the issues; the score is the share of seeded issues it names. 39 verified mechanically, 4 judgement items listed as such.
3 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.
Every skill tested for SEO, ranked.
Skill cohort (launch cells)
one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.
- Helped on someBest result in this test
skills/seo-onpage rampstackco/claude-skills*
Helped on Gemini 3.1 Flash Lite; no reliable difference on Claude Haiku 4.5 and GPT-5 mini. Measured against no skill loaded.
Found 56% of the seeded on-page issues, against 39% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · unclear. The model scored 24% without it.
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · unclear. The model scored 45% without it.
Adds about 6,200 tokens of context to every request.
- No reliable difference
skills/seo affaan-m/ECC
No reliable difference on any model tested. The models scored 39% without it. Measured against no skill loaded.
Found 44% of the seeded on-page issues, against 39% with no skill loaded. Instruction arm: not measured.
- Claude Haiku 4.5 · unclear. The model scored 24% without it.
- Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
- GPT-5 mini · unclear. The model scored 45% without it.
Adds about 1,600 tokens of context to every request.
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.
- No reliable differenceBest result in this test
skills/seo-onpage rampstackco/claude-skills*
No reliable difference on any model tested. The models scored 53% without it. Measured against the same task with a one-line instruction. Based on 60 of 68 runs; the rest returned nothing.
Found 63% of the seeded on-page issues, against 53% with no skill loaded. Against 53% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
- GPT-5 mini · unclear. The model scored 58% without it.
Adds about 6,000 tokens of context to every request.
- No reliable difference
marketing-skill/skills/seo-audit alirezarezvani/claude-skills
No reliable difference on any model tested. The models scored 47% without it. Measured against the same task with a one-line instruction.
Found 58% of the seeded on-page issues, against 47% with no skill loaded. Against 56% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
- GPT-5 mini · unclear. The model scored 46% without it.
Adds about 5,600 tokens of context to every request.
- No reliable difference
skills/seo affaan-m/ECC
No reliable difference on any model tested. The models scored 51% without it. Measured against the same task with a one-line instruction.
Found 56% of the seeded on-page issues, against 51% with no skill loaded. Against 56% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 49% without it.
- GPT-5 mini · unclear. The model scored 54% without it.
Adds about 1,600 tokens of context to every request.
Show the numbers
| Skill | Cohort | Without the skill | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|---|
| No skill loaded First run, with and without the skill | baseline | 39% | 39% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead First run, with and without the skill | baseline | instruction arm: not measured | 0 by definition | 0 | 2026-08-30 | ||
skills/seo-onpage from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this test | rampstackco/claude-skills* | 39% | 39% -> 56% | +0.17 [+0.04, +0.23] | In free drift on 2 of 3 | 6,184 | 2026-08-30 |
skills/seo from affaan-m/ECC d8409a4b0813 First run, with and without the skill | affaan-m/ECC | 39% | 39% -> 44% | +0.05 [+0.00, +0.12] | In free drift | 1,637 | 2026-08-30 |
| No skill loaded Wider run, three ways | baseline | 51% | 51% | 0 by definition | 0 | 2026-09-05 | |
| One-line instruction instead Wider run, three ways | baseline | 55% | 0 by definition | 0 | 2026-09-05 | ||
skills/seo-onpage from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this test | rampstackco/claude-skills* | 53% | 53% -> 63% | +0.10 [+0.08, +0.13] | In free drift | 5,975 | 2026-09-05 |
marketing-skill/skills/seo-audit from alirezarezvani/claude-skills 19392f7a0826 Wider run, three ways | alirezarezvani/claude-skills | 47% | 47% -> 58% | +0.02 [-0.04, +0.07] | In free drift | 5,625 | 2026-09-05 |
skills/seo from affaan-m/ECC d8409a4b0813 Wider run, three ways | affaan-m/ECC | 51% | 51% -> 56% | -0.00 [-0.06, +0.05] | In free drift | 1,578 | 2026-09-05 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | |||||||
In Claude Code
These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.
Skill cohort (one pass)
one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.
skills/seo-onpage
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
skills/seo
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
Panel v2 (repeat sampling)
two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.
skills/seo-onpage
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
skills/seo
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
Fable 5.1, one pass
one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.
skills/seo-onpage
- Claude Fable 5.1 · unclear in the version re-test
skills/seo
- Claude Fable 5.1 · unclear in the version re-test
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
skills/seo-onpage
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · unclear in the version re-test
skills/seo
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · unclear in the version re-test
marketing-skill/skills/seo-audit
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
| Skill | Claude Haiku 4.5 · skill-cohort · | Gemini 3.1 Flash Lite · skill-cohort · | GPT-5 mini · skill-cohort · | Gemini 3.1 Flash Lite · expansion-cohort · | GPT-5 mini · expansion-cohort · |
|---|---|---|---|---|---|
| rampstackco-seo-onpage · First run, with and without the skill | unclear · +0.235 | helped · +0.219 | unclear · +0.041 | Not run | Not run |
| rampstackco-seo-onpage · Wider run, three ways | Not run | Not run | Not run | unclear · +0.126 | unclear · +0.081 |
| affaan-m-seo · First run, with and without the skill | unclear · +0.000 | unclear · +0.122 | unclear · +0.032 | Not run | Not run |
| alirezarezvani-seo-audit · Wider run, three ways | Not run | Not run | Not run | unclear · +0.070 | unclear · -0.036 |
| affaan-m-seo · Wider run, three ways | Not run | Not run | Not run | unclear · +0.053 | unclear · -0.058 |
Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/seo at d8409a4b0813.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/seo from affaan-m/ECC | d8409a4b0813 | 9a655a52cfd9 | yes | Names auditing and on-page optimization in the same sentence. |
- obra/superpowers has no skill naming search or on-page work in any when-to-use.
- anthropics/skills has none either.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in alirezarezvani/claude-skills marketing-skill/skills/seo-audit at 19392f7a0826.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| marketing-skill/skills/seo-audit from alirezarezvani/claude-skills | 19392f7a0826 | 979f942727c2 | yes | Names the class label verbatim: "on-page SEO" and "audit" in the same sentence. The most literal match to a class name anywhere in the pass. Pinned. |
| marketing-skill/skills/programmatic-seo from alirezarezvani/claude-skills | 19392f7a0826 | 0851f7c8ab34 | no | Rules itself out in its own words and hands the class to seo-audit, the same shape as affaan-m/ECC skills/product-lens in the committed record. |
| marketing-skill/skills/local-seo-manager from alirezarezvani/claude-skills | 19392f7a0826 | 6d93fc13f7e0 | no | Local listings and map presence. Names a channel, not an on-page audit, and hands national SEO to seo-audit in its own words. |
- mattpocock/skills has no skill whose when-to-use names SEO or on-page work.
- addyosmani/agent-skills has none.
- openclaw/openclaw has none.
- wshobson/agents has none.
- github/awesome-copilot has none.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- Jeffallan/claude-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- yusufkaraaslan/Skill_Seekers has none.
Voice
The model is given a brand brief and asked for a short piece of copy. There is no answer key; a grader sees the with and without outputs, blind and in random order, and picks one. The score is how often the skill's output was preferred.
3 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.
Every skill tested for brand voice, ranked.
Skill cohort (launch cells)
one pass over each committed item, at temperature 0 where the vendor accepted it. Matrix hash bd0f7d55. No pre-registration: this run predates the practice on this arm. Run report.
- Helped clearlyBest result in this test
skills/brand-voice affaan-m/ECC
Helped clearly on Claude Haiku 4.5, Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.
A judge preferred its output 24 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.
- Claude Haiku 4.5 · helped
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · helped
Adds about 1,200 tokens of context to every request.
- Helped on some
skills/brand-voice rampstackco/claude-skills*
Helped on Gemini 3.1 Flash Lite; no reliable difference on Claude Haiku 4.5 and GPT-5 mini. Measured against no skill loaded.
A judge preferred its output 21 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.
- Claude Haiku 4.5 · unclear. The model scored 30% without it.
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · unclear. The model scored 40% without it.
Adds about 5,700 tokens of context to every request.
Show the numbers
| Skill | Cohort | Without the skill | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|---|
| No skill loaded | baseline | 25% | 25% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead | baseline | instruction arm: not applicable to a blind preference | 0 by definition | 0 | 2026-08-30 | ||
skills/brand-voice from affaan-m/ECC d8409a4b0813 Best result in this test | affaan-m/ECC | 20% | 20% -> 80% baseline shown as the complement | +0.60 [+0.60, +0.60] | Stable | 1,170 | 2026-08-30 |
skills/brand-voice from rampstackco/claude-skills 047924252254 | rampstackco/claude-skills* | 30% | 30% -> 70% baseline shown as the complement | +0.40 [+0.20, +0.60] | In free drift on 2 of 3 | 5,674 | 2026-08-30 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | |||||||
| Skill | Claude Haiku 4.5 · skill-cohort · | Gemini 3.1 Flash Lite · skill-cohort · | GPT-5 mini · skill-cohort · |
|---|---|---|---|
| affaan-m-brand-voice | helped · +0.600 | helped · +0.600 | helped · +0.600 |
| rampstackco-brand-voice | unclear · +0.400 | helped · +0.600 | unclear · +0.200 |
Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/brand-voice at d8409a4b0813.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/brand-voice from affaan-m/ECC | d8409a4b0813 | eae455eed766 | yes | Names writing voice and consistency of it. |
| skills/brand-guidelines from anthropics/skills | 3b3fad96af16 | 1120b3769e29 | no | Visual identity, colors and type. Brand is shared and voice is not. |
- obra/superpowers has no skill naming writing voice in any when-to-use.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in samber/cc-skills skills/copywriting-tone-of-voice-creator at 62bbac4f2f0c.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/copywriting-tone-of-voice-creator from samber/cc-skills | 62bbac4f2f0c | f4753aa281d4 | yes | Names defining brand voice in its when-to-use. It declares a fixed artifact, TONE.md, and a fixed section inventory, which is a closer output-format fit than the incumbent's own description states. Pinned. |
| marketing-skill/skills/content-humanizer from alirezarezvani/claude-skills | 19392f7a0826 | 394c104c2199 | no | The class named is removing AI tells from finished copy, and the skill excludes content creation in its own words. "Voice" appears once, inside a trigger list. Removing a generic register is not building or holding a specific one. |
| skills/finnish-humanizer from github/awesome-copilot | c956566a35c3 | a21c9d7e4fb7 | no | The same shape as content-humanizer, narrowed further to one natural language. QUOTATION CORRECTED AGAINST THE BYTES: recon section 2.5 ends this description at a full stop where the committed bytes carry a comma and continue for two further sentences, so the report quotes a truncation. It is recorded here as the artifact writes it. |
- mattpocock/skills has no skill whose when-to-use names writing voice.
- addyosmani/agent-skills has none.
- openclaw/openclaw has none.
- wshobson/agents has none.
- phuryn/pm-skills has none.
- Jeffallan/claude-skills has none.
- open-gsd/gsd-core has none.
- EveryInc/compound-engineering-plugin has none.
- yusufkaraaslan/Skill_Seekers has none.
Code review
10 fixture source files in 4 languages, each seeded with known bugs, security issues and correctness problems, 54 in total, drawn from a closed list of 18 defect types that both arms are given. The model is asked to name the symbol each defect sits on, never a line number; the score is the share of seeded defects it names. 50 are verified in the fixtures by mechanical check, 4 are judgement items listed as such, and 30 pieces of correct code that read as suspicious are recorded as traps and are deliberately not in the key.
8 repos have a qualifying skill in this class; 0 declared empty; 0 not yet verified. Which repositories, and what was found.
Every skill tested for code review, ranked.
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 1ee5d85f. Pre-registration · Run report.
- No reliable differenceBest result in this test
plugins/developer-essentials/skills/code-review-excellence wshobson/agents
No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction.
Found 63% of the seeded code defects, against 64% with no skill loaded. Against 62% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
- GPT-5 mini · unclear. The model scored 71% without it.
Adds about 3,900 tokens of context to every request.
- No reliable difference
skills/code-reviewer Jeffallan/claude-skills
No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.
Found 64% of the seeded code defects, against 63% with no skill loaded. Against 64% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
- GPT-5 mini · unclear. The model scored 69% without it.
Adds about 8,300 tokens of context to every request.
- No reliable difference
engineering/skills/pr-review-expert alirezarezvani/claude-skills
No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.
Found 60% of the seeded code defects, against 63% with no skill loaded. Against 65% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
- GPT-5 mini · unclear. The model scored 69% without it.
Adds about 4,300 tokens of context to every request.
- No reliable difference
skills/code-review-and-quality addyosmani/agent-skills
No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction.
Found 61% of the seeded code defects, against 64% with no skill loaded. Against 66% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
- GPT-5 mini · unclear. The model scored 71% without it.
Adds about 5,200 tokens of context to every request.
- No reliable difference
skills/gsd-code-review open-gsd/gsd-core
No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.
Found 59% of the seeded code defects, against 63% with no skill loaded. Against 64% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
- GPT-5 mini · unclear. The model scored 68% without it.
Adds about 1,400 tokens of context to every request.
- No reliable difference
skills/code-review-web rampstackco/claude-skills*
No reliable difference on any model tested. The models scored 64% without it. Measured against the same task with a one-line instruction.
Found 58% of the seeded code defects, against 64% with no skill loaded. Against 65% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
- GPT-5 mini · unclear. The model scored 71% without it.
Adds about 8,200 tokens of context to every request.
- No reliable difference
skills/engineering/code-review mattpocock/skills
No reliable difference on any model tested. The models scored 63% without it. Measured against the same task with a one-line instruction.
Found 57% of the seeded code defects, against 63% with no skill loaded. Against 69% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · unclear. The model scored 57% without it.
- GPT-5 mini · unclear. The model scored 69% without it.
Adds about 2,300 tokens of context to every request.
- Made worse
skills/ce-code-review EveryInc/compound-engineering-plugin
Made results worse on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction.
Found 34% of the seeded code defects, against 64% with no skill loaded. Against 64% with a one-line instruction instead.
- Gemini 3.1 Flash Lite · worse
- GPT-5 mini · unclear. The model scored 71% without it.
Adds about 73,700 tokens of context to every request.
Show the numbers
| Skill | Cohort | Without the skill | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|---|
| No skill loaded | baseline | 64% | 64% | 0 by definition | 0 | 2026-09-05 | |
| One-line instruction instead | baseline | 65% | 0 by definition | 0 | 2026-09-05 | ||
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents 38e19c20d2b1 Best result in this test | wshobson/agents | 64% | 64% -> 63% | +0.01 [+0.01, +0.02] | In free drift | 3,938 | 2026-09-05 |
skills/code-reviewer from Jeffallan/claude-skills 882ef55e377d | Jeffallan/claude-skills | 63% | 63% -> 64% | -0.00 [-0.05, +0.05] | In free drift | 8,290 | 2026-09-05 |
engineering/skills/pr-review-expert from alirezarezvani/claude-skills 19392f7a0826 | alirezarezvani/claude-skills | 63% | 63% -> 60% | -0.05 [-0.08, -0.01] | In free drift | 4,260 | 2026-09-05 |
skills/code-review-and-quality from addyosmani/agent-skills d2c37ef6225d | addyosmani/agent-skills | 64% | 64% -> 61% | -0.05 [-0.05, -0.04] | In free drift | 5,218 | 2026-09-05 |
skills/gsd-code-review from open-gsd/gsd-core 6beaa66b2587 | open-gsd/gsd-core | 63% | 63% -> 59% | -0.05 [-0.07, -0.03] | In free drift | 1,420 | 2026-09-05 |
skills/code-review-web from rampstackco/claude-skills a67dd34c609f | rampstackco/claude-skills* | 64% | 64% -> 58% | -0.07 [-0.14, -0.01] | In free drift | 8,206 | 2026-09-05 |
skills/engineering/code-review from mattpocock/skills 6654f6b60cd9 | mattpocock/skills | 63% | 63% -> 57% | -0.12 [-0.14, -0.10] | In free drift | 2,278 | 2026-09-05 |
skills/ce-code-review from EveryInc/compound-engineering-plugin c9c10f8c7541 | EveryInc/compound-engineering-plugin | 64% | 64% -> 34% | -0.30 [-0.59, -0.02] | Mixed, see page | 73,715 | 2026-09-05 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | |||||||
In Claude Code
These are tested in Claude Code, a second instrument. No figure here is averaged with one tested via API above, and this group is not ranked: it reports. How the two were compared.
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
skills/code-review-web
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
skills/engineering/code-review
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
engineering/skills/pr-review-expert
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
skills/code-review-and-quality
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
skills/code-reviewer
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
skills/ce-code-review
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
plugins/developer-essentials/skills/code-review-excellence
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
skills/gsd-code-review
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
Each cell is one measured pair: the verdict and the delta that cell holds. Nothing on this grid is averaged, across models, instruments or runs. Sorting orders rows by a column's delta; cells with no delta sort last in both directions. Claude Code columns name their run: the panel measured these models more than once and the runs are never pooled.
How this external slot was filled
Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section 3 "Identify the standards sources", in mattpocock/skills skills/engineering/code-review at 6654f6b60cd9.
THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart. THE CANDIDATE LIST HOLDS THE SEVEN THAT QUALIFIED, AND THE NON-QUALIFYING SKILLS ARE IN `emptyRepos` RATHER THAN AS CANDIDATES. Every candidate in this record carries committed bytes, which is what makes the losing side of a decision checkable; a candidate whose bytes cannot be committed would be a quoted sentence with nothing behind it. CherryHQ/cherry-studio .agents/skills/gh-pr-review names the class and is NOT recorded as a candidate for exactly that reason: the repository is AGPL-3.0 at its root, pinning its bytes into this repository would impose AGPL obligations on the surrounding work, and recon section 5.1 bars it. It is named here instead, with its reason, so the exclusion is auditable rather than silent.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/engineering/code-review from mattpocock/skills | 6654f6b60cd9 | dcded7969b1d | yes | Names the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned. |
| engineering/skills/pr-review-expert from alirezarezvani/claude-skills | 19392f7a0826 | dcded7969b1d | yes | Names the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned. |
| skills/code-review-and-quality from addyosmani/agent-skills | d2c37ef6225d | dcded7969b1d | yes | Names the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned. |
| skills/code-reviewer from Jeffallan/claude-skills | 882ef55e377d | dcded7969b1d | yes | Names the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned. |
| skills/ce-code-review from EveryInc/compound-engineering-plugin | c9c10f8c7541 | dcded7969b1d | yes | Names the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned. |
| plugins/developer-essentials/skills/code-review-excellence from wshobson/agents | 38e19c20d2b1 | dcded7969b1d | yes | Names the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned. |
| skills/gsd-code-review from open-gsd/gsd-core | 6beaa66b2587 | dcded7969b1d | yes | Names the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned. |
- openclaw/openclaw has no skill whose when-to-use names reviewing code.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- yusufkaraaslan/Skill_Seekers has none.
- obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
- affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
- anthropics/skills has no skill whose when-to-use names reviewing code.
- github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled
Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in alirezarezvani/claude-skills engineering/skills/pr-review-expert at 19392f7a0826.
THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/engineering/code-review from mattpocock/skills | 6654f6b60cd9 | 0b3ac983ef14 | yes | Names the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned. |
| engineering/skills/pr-review-expert from alirezarezvani/claude-skills | 19392f7a0826 | 0b3ac983ef14 | yes | Names the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned. |
| skills/code-review-and-quality from addyosmani/agent-skills | d2c37ef6225d | 0b3ac983ef14 | yes | Names the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned. |
| skills/code-reviewer from Jeffallan/claude-skills | 882ef55e377d | 0b3ac983ef14 | yes | Names the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned. |
| skills/ce-code-review from EveryInc/compound-engineering-plugin | c9c10f8c7541 | 0b3ac983ef14 | yes | Names the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned. |
| plugins/developer-essentials/skills/code-review-excellence from wshobson/agents | 38e19c20d2b1 | 0b3ac983ef14 | yes | Names the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned. |
| skills/gsd-code-review from open-gsd/gsd-core | 6beaa66b2587 | 0b3ac983ef14 | yes | Names the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned. |
- openclaw/openclaw has no skill whose when-to-use names reviewing code.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- yusufkaraaslan/Skill_Seekers has none.
- obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
- affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
- anthropics/skills has no skill whose when-to-use names reviewing code.
- github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in addyosmani/agent-skills skills/code-review-and-quality at d2c37ef6225d.
THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/engineering/code-review from mattpocock/skills | 6654f6b60cd9 | 2e6a37430eac | yes | Names the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned. |
| engineering/skills/pr-review-expert from alirezarezvani/claude-skills | 19392f7a0826 | 2e6a37430eac | yes | Names the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned. |
| skills/code-review-and-quality from addyosmani/agent-skills | d2c37ef6225d | 2e6a37430eac | yes | Names the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned. |
| skills/code-reviewer from Jeffallan/claude-skills | 882ef55e377d | 2e6a37430eac | yes | Names the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned. |
| skills/ce-code-review from EveryInc/compound-engineering-plugin | c9c10f8c7541 | 2e6a37430eac | yes | Names the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned. |
| plugins/developer-essentials/skills/code-review-excellence from wshobson/agents | 38e19c20d2b1 | 2e6a37430eac | yes | Names the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned. |
| skills/gsd-code-review from open-gsd/gsd-core | 6beaa66b2587 | 2e6a37430eac | yes | Names the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned. |
- openclaw/openclaw has no skill whose when-to-use names reviewing code.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- yusufkaraaslan/Skill_Seekers has none.
- obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
- affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
- anthropics/skills has no skill whose when-to-use names reviewing code.
- github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in Jeffallan/claude-skills skills/code-reviewer at 882ef55e377d.
THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/engineering/code-review from mattpocock/skills | 6654f6b60cd9 | f0f24eedceee | yes | Names the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned. |
| engineering/skills/pr-review-expert from alirezarezvani/claude-skills | 19392f7a0826 | f0f24eedceee | yes | Names the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned. |
| skills/code-review-and-quality from addyosmani/agent-skills | d2c37ef6225d | f0f24eedceee | yes | Names the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned. |
| skills/code-reviewer from Jeffallan/claude-skills | 882ef55e377d | f0f24eedceee | yes | Names the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned. |
| skills/ce-code-review from EveryInc/compound-engineering-plugin | c9c10f8c7541 | f0f24eedceee | yes | Names the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned. |
| plugins/developer-essentials/skills/code-review-excellence from wshobson/agents | 38e19c20d2b1 | f0f24eedceee | yes | Names the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned. |
| skills/gsd-code-review from open-gsd/gsd-core | 6beaa66b2587 | f0f24eedceee | yes | Names the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned. |
- openclaw/openclaw has no skill whose when-to-use names reviewing code.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- yusufkaraaslan/Skill_Seekers has none.
- obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
- affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
- anthropics/skills has no skill whose when-to-use names reviewing code.
- github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in EveryInc/compound-engineering-plugin skills/ce-code-review at c9c10f8c7541.
THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/engineering/code-review from mattpocock/skills | 6654f6b60cd9 | bd5c9fac33fb | yes | Names the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned. |
| engineering/skills/pr-review-expert from alirezarezvani/claude-skills | 19392f7a0826 | bd5c9fac33fb | yes | Names the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned. |
| skills/code-review-and-quality from addyosmani/agent-skills | d2c37ef6225d | bd5c9fac33fb | yes | Names the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned. |
| skills/code-reviewer from Jeffallan/claude-skills | 882ef55e377d | bd5c9fac33fb | yes | Names the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned. |
| skills/ce-code-review from EveryInc/compound-engineering-plugin | c9c10f8c7541 | bd5c9fac33fb | yes | Names the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned. |
| plugins/developer-essentials/skills/code-review-excellence from wshobson/agents | 38e19c20d2b1 | bd5c9fac33fb | yes | Names the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned. |
| skills/gsd-code-review from open-gsd/gsd-core | 6beaa66b2587 | bd5c9fac33fb | yes | Names the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned. |
- openclaw/openclaw has no skill whose when-to-use names reviewing code.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- yusufkaraaslan/Skill_Seekers has none.
- obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
- affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
- anthropics/skills has no skill whose when-to-use names reviewing code.
- github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in wshobson/agents plugins/developer-essentials/skills/code-review-excellence at 38e19c20d2b1.
THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/engineering/code-review from mattpocock/skills | 6654f6b60cd9 | d796f1a58081 | yes | Names the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned. |
| engineering/skills/pr-review-expert from alirezarezvani/claude-skills | 19392f7a0826 | d796f1a58081 | yes | Names the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned. |
| skills/code-review-and-quality from addyosmani/agent-skills | d2c37ef6225d | d796f1a58081 | yes | Names the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned. |
| skills/code-reviewer from Jeffallan/claude-skills | 882ef55e377d | d796f1a58081 | yes | Names the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned. |
| skills/ce-code-review from EveryInc/compound-engineering-plugin | c9c10f8c7541 | d796f1a58081 | yes | Names the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned. |
| plugins/developer-essentials/skills/code-review-excellence from wshobson/agents | 38e19c20d2b1 | d796f1a58081 | yes | Names the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned. |
| skills/gsd-code-review from open-gsd/gsd-core | 6beaa66b2587 | d796f1a58081 | yes | Names the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned. |
- openclaw/openclaw has no skill whose when-to-use names reviewing code.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- yusufkaraaslan/Skill_Seekers has none.
- obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
- affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
- anthropics/skills has no skill whose when-to-use names reviewing code.
- github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
How this external slot was filled
Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in open-gsd/gsd-core skills/gsd-code-review at 6beaa66b2587.
THE COMMITTED BASIS, QUOTED FROM THE CLASS RECORD. The code-review class is scored by `scoreLedgerCoverage` over `harness/claims/claims-skills.json` S11-code-review-web, whose task template states to BOTH arms: "Report one finding per line, in exactly this form and nothing else:" / "<symbol> :: <code>" / "The symbol names the function, constant or named query the defect sits on, written as the file writes it" / "Do not report line numbers." / "No headings, no prose, no explanations, no summary." The basis has three components: a ROOT-CAUSE SYMBOL anchor, a CLOSED CODE vocabulary of 18 codes, and a flat per-finding record with no prose, scored against the seeded ledger at harness/fixtures/code-review/defect-ledger.json. HOW THE SEVEN SIT AGAINST IT, AS ORDERING AND NOT AS ADMISSION. `mattpocock/skills` code-review fixes the only CLOSED CODE VOCABULARY of the seven: twelve named Fowler smells carried as a baseline that applies even when the repo documents nothing, named per finding. `alirezarezvani/claude-skills` pr-review-expert fixes the only FLAT PER-FINDING RECORD in its own bytes, a literal Output Format block with numbered findings under fixed severity bands, and anchors it on `file:line`, which the class template rules out in as many words. NEITHER SATISFIES THE SYMBOL COMPONENT and no candidate does: not one anchors a finding on the function, constant or named query the defect sits on. The closed-code component cannot separate them either, because the class hands its own 18 codes to BOTH arms, so a skill bringing a vocabulary of its own is not advantaged by the basis. `addyosmani/agent-skills` code-review-and-quality and `Jeffallan/claude-skills` code-reviewer prescribe a sectioned narrative report with a mandatory summary and verdict, and Jeffallan's frontmatter declares `output-format: report`; that is the prose the class template forbids, and it is a statement about how far their output sits from the scored one, not a reason to leave their bytes uncommitted. `EveryInc/compound-engineering-plugin` ce-code-review and `open-gsd/gsd-core` gsd-code-review fix no output format in their pinned bytes at all, deferring it to a reference file and a repo config, and to a REVIEW.md in a GSD phase directory. `wshobson/agents` code-review-excellence prescribes no output format, being about review practice and mentoring. THE RECON REFUSED TO RANK THESE SEVEN and said so: with no committed output format only `names-the-class` applied, and its Finding 7 recorded that the three it listed were separated from the other four by star count alone, which is not a criterion in the rule. This is the re-run that finding asked for, against the basis PR #38 committed with the class. Under the cap it would have produced a tie between the first two and left five unpinned. Under R23 all seven are pinned and the analysis is what a reader uses to tell them apart.
| Candidate | Commit | Content hash | Qualified | Why, in our words |
|---|---|---|---|---|
| skills/engineering/code-review from mattpocock/skills | 6654f6b60cd9 | 9d5072bfb96a | yes | Names the class. Fixes the only closed code vocabulary of the seven, twelve named Fowler smells applied per finding, and fixes no symbol anchor. Pinned. |
| engineering/skills/pr-review-expert from alirezarezvani/claude-skills | 19392f7a0826 | 9d5072bfb96a | yes | Names the class. Fixes the only flat per-finding record in its own bytes, and anchors it on file:line, which the class template rules out. Pinned. |
| skills/code-review-and-quality from addyosmani/agent-skills | d2c37ef6225d | 9d5072bfb96a | yes | Names the class. Its template is a sectioned narrative report with Context, five axis sections, Verification and Verdict, and fixes no per-finding shape; its severity prefixes rank a finding rather than naming one. Pinned. |
| skills/code-reviewer from Jeffallan/claude-skills | 882ef55e377d | 9d5072bfb96a | yes | Names the class. Frontmatter declares `output-format: report` and the Output Template requires a summary, praise, questions for the author and a verdict, which is the prose the class template forbids. Pinned. |
| skills/ce-code-review from EveryInc/compound-engineering-plugin | c9c10f8c7541 | 9d5072bfb96a | yes | Names the class. The pinned SKILL.md fixes no output format, deferring it to references/modes-and-output.md and to a repo-level .compound-engineering/config.yaml. Pinned. |
| plugins/developer-essentials/skills/code-review-excellence from wshobson/agents | 38e19c20d2b1 | 9d5072bfb96a | yes | Names the class. Prescribes no output format at all, being about review practice, standards-setting and mentoring rather than about performing a review. Pinned. |
| skills/gsd-code-review from open-gsd/gsd-core | 6beaa66b2587 | 9d5072bfb96a | yes | Names the class. Its output is a REVIEW.md artifact in a GSD phase directory and its execution context lives outside the skill directory, so the pinned bytes fix no format. Pinned. |
- openclaw/openclaw has no skill whose when-to-use names reviewing code.
- phuryn/pm-skills has none.
- samber/cc-skills has none.
- yusufkaraaslan/Skill_Seekers has none.
- obra/superpowers skills/requesting-code-review dispatches a reviewer and does not perform a review, and skills/receiving-code-review is downstream of the class.
- affaan-m/ECC skills/flutter-dart-code-review names the class and binds it to Flutter/Dart and one of six named state-management libraries, so it is not a language-general code review.
- anthropics/skills has no skill whose when-to-use names reviewing code.
- github/awesome-copilot skills/postgresql-code-review and skills/sql-code-review are dialect-specific review of SQL rather than general code review.
Test writing
10 small modules in 2 languages, pure functions and small classes with no input or output of their own, each with a committed list of behaviours it is claimed to obey, 50 in total. The model is asked to write tests for the module and is never shown the list. A behaviour counts as covered only when the model's tests pass against the correct module and fail against a committed broken copy of it that changes exactly that one behaviour; the score is the share of behaviours covered. A test file that fails against the correct module scores zero for that module and the reason is recorded. 20 tests that read as though they cover a behaviour and would catch nothing are committed as traps and are deliberately not in the key.
0 repos have a qualifying skill in this class; 3 declared empty; 12 not yet verified. Which repositories, and what was found.
The instrument for this class is committed and no model has been called against it, so there is no table here. The fixtures, the answer key and the task set are in the repository and the projected size of the run is recorded beside them; an effect, an interval and a cost appear when the run happens and not before.
THE OWN SLOT FOR THIS CLASS IS EMPTY AND IS SAID TO BE EMPTY. THIS IS THE FIRST OWN SLOT IN THE COHORT THAT IS, and the reason is that the operator has no skill for this class. Every own slot before this one existed because the class was built around a skill the operator had already published; this class was built first and then looked for one. The repository at skills/ was read at the commit pinned on the code review slot, a67dd34c609f034c0cfd736a348659bbdf1605bf, and the two nearest candidates were weighed by the same standard the selection rule applies to an external slot, which is whether the when-to-use NAMES the class. skills/qa-testing names running QA on a page, a feature or a site at three depth tiers, after a deploy or before a launch, and hands code-level work to code-review-web in its own When NOT to use; it is verification of a running site, not the writing of a test for a module. skills/code-review-web names reviewing a pull request, debugging a production issue, investigating a build failure and auditing security, and hands pre-launch QA to qa-testing; it is review of code, not the writing of tests for it. No other skill in the repository names writing tests in any when-to-use: skills/frontend-component-build lists unit tests once as one row of a testing checklist for a component it is building, which is a step inside a different class rather than this one. Neither candidate is pinned, because a slot filled with the nearest thing to hand would make every number in this class a measurement of something other than what the class is named for.
Read for a skill naming this class, and none found: rampstackco/claude-skills. This class publishes NO row at all rather than one or a pair. It holds no claim, so nothing is planned, nothing is projected as a call against an arm, and no cell, effect, interval or verdict exists for it. What is committed is the instrument: the fixtures, the answer key, the runner and the task set. If the operator publishes a skill for this class, this slot is filled in its own change on the same footing as any other slot, and the external slot below is selected then.
THE EXTERNAL SLOT FOR THIS CLASS IS EMPTY AND IS SAID TO BE EMPTY. The selection rule is unchanged and is the one recorded with the cohort. It has not been run against these repositories for this class, so no candidate has been weighed, none is pinned and none is named here: naming a likely candidate before the rule had been applied to it would be a guess dressed as a record. The rule is run in its own change against a source list verified at the commit it names, and the candidates, their quoted sentences, their commits and the decision are recorded there on the same footing as every other external slot.
Repositories in scope, from the same selection rule as every other external slot: obra/superpowers, affaan-m/ECC, anthropics/skills. With the own slot for this class also empty, the class has no arm on either side. That is a different thing from a class with one arm and it is recorded as two entries rather than one, because the two are empty for different reasons and only one of them has had the rule applied to it.