Claude skills for Writing Skills, tested

Skills for writing skills work, measured on skill authoring tasks. Each was run on the same task pairs with the skill loaded and without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.

Measured on 3 models (Claude, GPT & Gemini): Claude Haiku 4.5, GPT-5 mini, Gemini 3.1 Flash Lite.

The model is given a short brief and asked to write a complete SKILL.md. The score is the share of the open Agent Skills specification's requirements the result meets, counting only the 14 defined requirements that apply to the document written; requirements for optional fields the document does not use are not counted. A second reading against this site's own house style is shown on each skill's page but is not ranked, because one of the skills tested is ours.

What helped, and on how many models

  • Tied for first · shared with 1 other, and the tie is broken by nothing

    skills/skill-creation-walkthrough rampstackco/claude-skills*

    Skill authoring

    Helped clearly on Claude Haiku 4.5 and GPT-5 mini; unclear on Gemini 3.1 Flash Lite. Measured against no skill loaded.

    Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · helped
    • GPT-5 mini · helped
    • Gemini 3.1 Flash Lite · unclear

    In Claude Code

    • Claude Fable 5, Skill cohort (one pass) · unclear
    • Claude Opus 5, Skill cohort (one pass) · not measured
    • Claude Fable 5, Panel v2 (repeat sampling) · unclear
    • Claude Haiku 4.5, Panel v2 (repeat sampling) · helped
    • Claude Opus 5, Panel v2 (repeat sampling) · not measured
    • Claude Sonnet 5, Panel v2 (repeat sampling) · unclear
    • Claude Fable 5.1, Fable 5.1, one pass · not measured

    Cost per task, with the skill loaded: about $0.00826.Adds about 8,400 tokens of context to every request.

  • Tied for first · shared with 1 other, and the tie is broken by nothing

    skills/skill-creator anthropics/skills

    Skill authoring

    Helped clearly on Claude Haiku 4.5 and GPT-5 mini; unclear on Gemini 3.1 Flash Lite. Measured against no skill loaded.

    Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · helped
    • GPT-5 mini · helped
    • Gemini 3.1 Flash Lite · unclear

    In Claude Code

    • Claude Fable 5, Skill cohort (one pass) · not measured
    • Claude Opus 5, Skill cohort (one pass) · not measured
    • Claude Fable 5, Panel v2 (repeat sampling) · not measured
    • Claude Haiku 4.5, Panel v2 (repeat sampling) · helped
    • Claude Opus 5, Panel v2 (repeat sampling) · unclear
    • Claude Sonnet 5, Panel v2 (repeat sampling) · unclear
    • Claude Fable 5.1, Fable 5.1, one pass · worse

    Cost per task, with the skill loaded: about $0.0116.Adds about 11,500 tokens of context to every request.

The numbers

How to read these tables

How to read these tables

Without -> with
The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
Effect (range)
With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
Orbit
  • Stable: delta at or above the pass threshold, and the interval excludes zero
  • Past the horizon: delta at or below the failure floor, and the interval excludes zero
  • In free drift: the interval spans zero, or the delta sits between the floor and the threshold
  • Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound

A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.

Cohort
The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
Commit
The exact version tested. Later commits are not covered by this verdict.
Injected context tokens
How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
Tested
The date. Verdicts age as models change.

Skill authoring

SkillCohortWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loadedbaseline74%0 by definition02026-08-30
One-line instruction insteadbaselineinstruction arm: not measured0 by definition02026-08-30
skills/skill-creation-walkthrough from rampstackco/claude-skills 047924252254 Best result in this testrampstackco/claude-skills*74% -> 100%+0.26 [+0.13, +0.38]Stable on 2 of 38,3902026-08-30
skills/skill-creator from anthropics/skills 3b3fad96af16anthropics/skills74% -> 100%+0.26 [+0.13, +0.38]Stable on 2 of 311,5142026-08-30
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.