Skills, tested

Ten skills across five task classes. Each is run twice on the same task pairs: once with the skill loaded as system context, once with the identical prompt and nothing injected.

Skill authoring

Deterministic, scored against a committed answer key. Own and external skills are listed together and ranked only against each other.

SkillCohortCommitInjected context tokensEffectTested
No skill loadedbaselinenot applicable00 by constructionnot yet
skills/skill-creator from anthropics/skillsexternal3b3fad96af1610,587not yet measurednot yet
skills/skill-creation-walkthrough from rampstackco/claude-skills (operated by this site)own0479242522548,687not yet measurednot yet
How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in anthropics/skills skills/skill-creator at 3b3fad96af16.

Two candidates name the class. The task class is scored on a single SKILL.md whose frontmatter must parse and whose section order must match SKILL_AUTHORING.md. writing-skills prescribes that document's section order under the heading above. skill-creator prescribes an authoring and evaluation workflow, and the structure it prescribes under its own Report structure heading is an evaluation report rather than the SKILL.md. writing-skills is the closer output format.

Addendum. RECORDED AFTER THE FACT, ON 2026-08-26, AND NOT ACTED ON. This tiebreak was decided when the skill-authoring class was scored against SKILL_AUTHORING.md alone. That document is now Key B, disclosed and unranked, and the ranked key is the open Agent Skills specification. The tiebreak reasoning above therefore turns on a basis that is no longer the ranked one. It is left exactly as it was decided, because a selection record that is rewritten to agree with a later ruling is not a record of what was decided. The selection was NOT re-run against Key A, and whether it would produce the same candidate is an open question stated here rather than assumed.

Second addendum. SECOND ADDENDUM, 2026-08-26. THE TIEBREAK WAS RE-RUN WITH KEY A AS THE OUTPUT-FORMAT BASIS, AND IT SWAPPED. Basis: the selection rule is unchanged, but the output format the class is scored on is now the open Agent Skills specification rather than SKILL_AUTHORING.md. Candidates compared: the same two that qualified, obra/superpowers skills/writing-skills and anthropics/skills skills/skill-creator. Result: skill-creator. It prescribes the specification's own model of the artifact, naming the required and the optional frontmatter fields, the bundled directory layout, the progressive disclosure levels and the line ceiling. writing-skills prescribes a body section order, which is the one part of a SKILL.md the specification explicitly leaves unconstrained, and it disagrees with the specification in three places: it puts the thousand and twenty four character limit on the whole frontmatter rather than on the description, it tells an author the description must not say what the skill does where the specification asks for both, and the example name in its own template is capitalised where the specification allows lowercase only. Run through Key A, writing-skills' own template fails a predicate and skill-creator's passes every applicable one. The original record and the first addendum above are left exactly as written. obra/superpowers skills/writing-skills moves to the rejected pins, where its bytes stay committed and its quoted sentence stays resolvable.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/writing-skills from obra/superpowersb36e0829c6d0d34db5c8aed6yesNames the class in its when-to-use directly.
skills/skill-creator from anthropics/skills3b3fad96af16e51ce07ea7fdyesAlso names the class directly, so the tiebreak applies.
skills/skill-scout from affaan-m/ECCd8409a4b081371b1e8a4aedenoSearches for an existing skill before one is written. It names the moment before authoring, not authoring.
skills/skill-comply from affaan-m/ECCd8409a4b0813e26cbbc30a32noMeasures whether a written skill is obeyed. It is about a skill, not about writing one.

Accessibility audit

Deterministic, scored against a committed answer key. Own and external skills are listed together and ranked only against each other.

SkillCohortCommitInjected context tokensEffectTested
No skill loadedbaselinenot applicable00 by constructionnot yet
skills/accessibility from affaan-m/ECCexternald8409a4b08131,546not yet measurednot yet
skills/accessibility-audit from rampstackco/claude-skills (operated by this site)own04792425225410,239not yet measurednot yet
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/accessibility at d8409a4b0813.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/accessibility from affaan-m/ECCd8409a4b0813ab86a50717d6yesNames auditing against WCAG in the when-to-use itself.
skills/click-path-audit from affaan-m/ECCd8409a4b0813662fc5d02483noAn audit of behavioural state, not of accessibility. The word audit is shared and the class is not.
  • obra/superpowers has no skill naming accessibility in any when-to-use.
  • anthropics/skills has none either; webapp-testing names browser testing rather than accessibility.

Spec writing

Deterministic, scored against a committed answer key, and paired blind comparison, win rate with ties counting half. Reported as two rows, never averaged.. Own and external skills are listed together and ranked only against each other.

SkillCohortCommitInjected context tokensEffectTested
No skill loadedbaselinenot applicable00 by constructionnot yet
skills/product-capability from affaan-m/ECCexternald8409a4b08131,046not yet measurednot yet
skills/pm-spec-writing from rampstackco/claude-skills (operated by this site)own0479242522546,292not yet measurednot yet
How this external slot was filled

Decided by: two candidates named the class, so the closer output format took it. The deciding text is at: SKILL.md, section heading, in affaan-m/ECC skills/product-capability at d8409a4b0813.

Two candidates name the class. The task class is scored on a document with required sections and acceptance criteria, so the closer output format is the one that fixes its sections. product-capability declares a canonical artifact and a section-by-section output format under the heading above. doc-coauthoring prescribes a three-stage collaboration and fixes no sections in the document it produces.

Addendum. RECORDED AFTER THE FACT, ON 2026-08-26, AND NOT ACTED ON. This tiebreak was decided when the skill-authoring class was scored against SKILL_AUTHORING.md alone. That document is now Key B, disclosed and unranked, and the ranked key is the open Agent Skills specification. The tiebreak reasoning above therefore turns on a basis that is no longer the ranked one. It is left exactly as it was decided, because a selection record that is rewritten to agree with a later ruling is not a record of what was decided. The selection was NOT re-run against Key A, and whether it would produce the same candidate is an open question stated here rather than assumed.

Second addendum. SECOND ADDENDUM, 2026-08-26. RE-EXAMINED, NOT RE-RUN, AND THE CANDIDATE IS UNCHANGED. Basis: this tiebreak was never decided on the operator's house standard. The spec-writing class is scored on the section list and the acceptance-criterion pattern that the task prompt states to BOTH arms, which is a contract the two arms are given rather than a convention one of them was taught, and R7 did not move it. Key A is the answer key for the skill-authoring class and has no bearing here. Candidates compared: the same two that qualified, affaan-m/ECC skills/product-capability and anthropics/skills skills/doc-coauthoring. Result: unchanged, product-capability. THE FIRST ADDENDUM ABOVE IS WRONG ON THIS SLOT and is left in place rather than edited. It was applied to both output-format tiebreaks at once and says this one was decided against SKILL_AUTHORING.md, which it was not. Correcting it by rewriting would hide that the error was made; this sentence is the correction.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/product-capability from affaan-m/ECCd8409a4b08133e0052802969yesNames writing a specification from product intent.
skills/doc-coauthoring from anthropics/skills3b3fad96af162e47d78846fayesNames technical specs among several document kinds, so the tiebreak applies.
skills/writing-plans from obra/superpowersb36e0829c6d048508f44bbfdnoTakes a spec as its input. It is downstream of the class, not the class.
skills/product-lens from affaan-m/ECCd8409a4b0813d082be7c3dd9noRules itself out in its own words and hands the class to product-capability.

On-page audit

Deterministic, scored against a committed answer key. Own and external skills are listed together and ranked only against each other.

SkillCohortCommitInjected context tokensEffectTested
No skill loadedbaselinenot applicable00 by constructionnot yet
skills/seo from affaan-m/ECCexternald8409a4b08131,018not yet measurednot yet
skills/seo-onpage from rampstackco/claude-skills (operated by this site)own0479242522545,117not yet measurednot yet
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/seo at d8409a4b0813.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/seo from affaan-m/ECCd8409a4b08139a655a52cfd9yesNames auditing and on-page optimization in the same sentence.
  • obra/superpowers has no skill naming search or on-page work in any when-to-use.
  • anthropics/skills has none either.

Voice

Paired blind comparison, win rate with ties counting half. Own and external skills are listed together and ranked only against each other.

SkillCohortCommitInjected context tokensEffectTested
No skill loadedbaselinenot applicable00 by constructionnot yet
skills/brand-voice from affaan-m/ECCexternald8409a4b08131,122not yet measurednot yet
skills/brand-voice from rampstackco/claude-skills (operated by this site)own0479242522545,236not yet measurednot yet
How this external slot was filled

Decided by: one candidate named the class in its own when-to-use. The deciding text is at: SKILL.md frontmatter, description, in affaan-m/ECC skills/brand-voice at d8409a4b0813.

Candidates considered, all of them pinned
CandidateCommitContent hashQualifiedWhy, in our words
skills/brand-voice from affaan-m/ECCd8409a4b0813eae455eed766yesNames writing voice and consistency of it.
skills/brand-guidelines from anthropics/skills3b3fad96af161120b3769e29noVisual identity, colors and type. Brand is shared and voice is not.
  • obra/superpowers has no skill naming writing voice in any when-to-use.

Nothing has been measured yet

The cohort is pinned, the task sets are frozen, the answer keys are committed and the cost of the run is projected. No model has been called, so no effect, no interval and no cost appears above. The table shows the instrument, not a result.

Selection for the five external slots follows one rule, recorded at harness/claims/skill-cohort.json and applied on 2026-08-26: From obra/superpowers, affaan-m/ECC and anthropics/skills, choose the skill whose when-to-use or equivalent most directly names the task class. If two qualify, take the one with the closer output format. If none qualifies in a repo, the slot may go to a second candidate from another listed repo. If no listed repo has a match for a class, the slot is left empty and said to be empty. A skill appears here as an identity pin and a measurement. Nothing else about it is published.