Claude skills for Writing Skills, tested

Skills for writing skills work, measured on skill authoring tasks. Each was the same task, run twice, once with the skill loaded and once without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.

Measured on 3 models (Claude, GPT & Gemini): Claude Haiku 4.5, GPT-5 mini, Gemini 3.1 Flash Lite.

The model is given a short brief and asked to write a complete SKILL.md. The score is the share of the open Agent Skills specification's requirements the result meets, counting only the 14 defined requirements that apply to the document written; requirements for optional fields the document does not use are not counted. A second reading against this site's own house style is shown on each skill's page but is not ranked, because one of the skills tested is ours.

What helped, and on how many models

  • Tied for first · shared with 4 others, and the tie is broken by nothing

    skills/skill-creation-walkthrough rampstackco/claude-skills*

    Skill authoring

    Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction.

    Found 100% of the specification requirements, against 76% with no skill loaded. Against 68% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    In Claude Code

    • Claude Fable 5, First run, single attempt · unclear
    • Claude Opus 5, First run, single attempt · not measured
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · helped
    • Claude Opus 5, Second run, every task twice · not measured
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · not measured
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · helped
    • Claude Opus 5, Wider run, three ways · not measured
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00415. Adds about 8,200 tokens of context to every request.

  • Tied for first · shared with 4 others, and the tie is broken by nothing

    skills/skill-creator anthropics/skills

    Skill authoring

    Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction.

    Found 100% of the specification requirements, against 82% with no skill loaded. Against 70% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    In Claude Code

    • Claude Fable 5, First run, single attempt · not measured
    • Claude Opus 5, First run, single attempt · not measured
    • Claude Fable 5, Second run, every task twice · not measured
    • Claude Haiku 4.5, Second run, every task twice · helped
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · worse
    • Claude Fable 5.1, Wider run, three ways · worse
    • Claude Haiku 4.5, Wider run, three ways · helped
    • Claude Opus 5, Wider run, three ways · not measured
    • Claude Sonnet 5, Wider run, three ways · not measured

    Cost per task, with the skill loaded: about $0.0199. Adds about 10,900 tokens of context to every request.

  • Tied for first · shared with 4 others, and the tie is broken by nothing

    skills/skill-creation-walkthrough rampstackco/claude-skills*

    Skill authoring

    Helped clearly on Claude Haiku 4.5 and GPT-5 mini; helped a little on Gemini 3.1 Flash Lite, under our bar. Measured against no skill loaded.

    Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · helped

    In Claude Code

    • Claude Fable 5, First run, single attempt · unclear
    • Claude Opus 5, First run, single attempt · not measured
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · helped
    • Claude Opus 5, Second run, every task twice · not measured
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · not measured
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · helped
    • Claude Opus 5, Wider run, three ways · not measured
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00826. Adds about 8,400 tokens of context to every request.

  • Tied for first · shared with 4 others, and the tie is broken by nothing

    distribution/claude-plugin/skills/skill-builder yusufkaraaslan/Skill_Seekers

    Skill authoring

    Helped clearly on Gemini 3.1 Flash Lite and GPT-5 mini. Measured against the same task with a one-line instruction.

    Found 99% of the specification requirements, against 70% with no skill loaded. Against 73% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · not measured
    • Claude Haiku 4.5, Wider run, three ways · helped
    • Claude Opus 5, Wider run, three ways · not measured
    • Claude Sonnet 5, Wider run, three ways · not measured

    Cost per task, with the skill loaded: about $0.00288. Adds about 1,200 tokens of context to every request.

  • Tied for first · shared with 4 others, and the tie is broken by nothing

    skills/skill-creator anthropics/skills

    Skill authoring

    Helped clearly on Claude Haiku 4.5 and GPT-5 mini; helped a little on Gemini 3.1 Flash Lite, under our bar. Measured against no skill loaded.

    Found 100% of the specification requirements, against 74% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · helped

    In Claude Code

    • Claude Fable 5, First run, single attempt · not measured
    • Claude Opus 5, First run, single attempt · not measured
    • Claude Fable 5, Second run, every task twice · not measured
    • Claude Haiku 4.5, Second run, every task twice · helped
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · worse
    • Claude Fable 5.1, Wider run, three ways · worse
    • Claude Haiku 4.5, Wider run, three ways · helped
    • Claude Opus 5, Wider run, three ways · not measured
    • Claude Sonnet 5, Wider run, three ways · not measured

    Cost per task, with the skill loaded: about $0.0116. Adds about 11,500 tokens of context to every request.

  • Second

    skills/skill-creator openclaw/openclaw

    Skill authoring

    Helped on Gemini 3.1 Flash Lite; helped a little on GPT-5 mini, under our bar. Measured against the same task with a one-line instruction.

    Found 99% of the specification requirements, against 75% with no skill loaded. Against 77% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · not measured
    • Claude Haiku 4.5, Wider run, three ways · helped
    • Claude Opus 5, Wider run, three ways · not measured
    • Claude Sonnet 5, Wider run, three ways · not measured

    Cost per task, with the skill loaded: about $0.00457. Adds about 600 tokens of context to every request.

The numbers

How to read these tables

How to read these tables

Without -> with
The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
Effect (range)
With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
Orbit
  • Stable: delta at or above the pass threshold, and the interval excludes zero
  • Past the horizon: delta at or below the failure floor, and the interval excludes zero
  • In free drift: the interval spans zero, or the delta sits between the floor and the threshold
  • Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound

A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.

Cohort
The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
Commit
The exact version tested. Later commits are not covered by this verdict.
Injected context tokens
How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
Tested
The date. Verdicts age as models change.

Skill authoring

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

distribution/claude-plugin/skills/skill-builder Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creation-walkthrough Claude Haiku 4.5 First run, with and without the skill, Deterministic

Holds

No skill: 0.618. With the skill: 1.000. Change +0.382. It could sit between +0.287 and +0.477.

skills/skill-creation-walkthrough Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creation-walkthrough Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.873. With the skill: 1.000. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creation-walkthrough GPT-5 mini Wider run, three ways, Deterministic

Holds

No skill: 0.655. With the skill: 1.000. Instruction only: 0.618. Change +0.345. It could sit between +0.176 and +0.515.

skills/skill-creation-walkthrough GPT-5 mini First run, with and without the skill, Deterministic

Holds

No skill: 0.718. With the skill: 1.000. Change +0.282. It could sit between +0.128 and +0.435.

skills/skill-creator Claude Haiku 4.5 First run, with and without the skill, Deterministic

Holds

No skill: 0.618. With the skill: 1.000. Change +0.382. It could sit between +0.287 and +0.477.

skills/skill-creator Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creator Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Holds

No skill: 0.873. With the skill: 1.000. Instruction only: 0.745. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creator Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.873. With the skill: 1.000. Change +0.127. It could sit between +0.056 and +0.199.

skills/skill-creator GPT-5 mini Wider run, three ways, Deterministic

Holds

No skill: 0.773. With the skill: 1.000. Instruction only: 0.664. Change +0.227. It could sit between +0.105 and +0.350.

skills/skill-creator GPT-5 mini First run, with and without the skill, Deterministic

Holds

No skill: 0.727. With the skill: 1.000. Change +0.273. It could sit between +0.083 and +0.462.

skills/skill-creator GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.627. With the skill: 0.983. Instruction only: 0.791. Change +0.356. It could sit between +0.198 and +0.514.

distribution/claude-plugin/skills/skill-builder GPT-5 mini Wider run, three ways, Deterministic

Holds

No skill: 0.536. With the skill: 0.982. Instruction only: 0.709. Change +0.445. It could sit between +0.268 and +0.623.

In Claude Code tested in Claude Code

distribution/claude-plugin/skills/skill-builder Claude Fable 5.1 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

distribution/claude-plugin/skills/skill-builder Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.600. With the skill: 1.000. Change +0.400. It could sit between +0.293 and +0.507.

distribution/claude-plugin/skills/skill-builder Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

distribution/claude-plugin/skills/skill-builder Claude Sonnet 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.545. With the skill: 1.000. Change +0.455. It could sit between +0.455 and +0.455.

skills/skill-creation-walkthrough Claude Haiku 4.5 Second run, every task twice, Deterministic

Holds

No skill: 0.691. With the skill: 1.000. Change +0.309. It could sit between +0.220 and +0.398.

skills/skill-creation-walkthrough Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Opus 5 Second run, every task twice, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Opus 5 First run, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Fable 5 Second run, every task twice, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Fable 5 First run, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Fable 5.1 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.564. With the skill: 1.000. Change +0.436. It could sit between +0.353 and +0.520.

skills/skill-creator Claude Haiku 4.5 Wider run, three ways, Deterministic

Holds

No skill: 0.655. With the skill: 1.000. Change +0.345. It could sit between +0.237 and +0.454.

skills/skill-creator Claude Haiku 4.5 Second run, every task twice, Deterministic

Holds

No skill: 0.709. With the skill: 1.000. Change +0.291. It could sit between +0.227 and +0.355.

skills/skill-creator Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Opus 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Opus 5 First run, single attempt, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Sonnet 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creator Claude Sonnet 5 Wider run, three ways, Deterministic

Could not measure

No skill: 1.000. With the skill: 1.000. Change +0.000. It could sit between +0.000 and +0.000.

skills/skill-creation-walkthrough Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.927. Change -0.073. It could sit between -0.168 and +0.022.

skills/skill-creator Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.927. Change -0.073. It could sit between -0.168 and +0.022.

skills/skill-creation-walkthrough Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.927. Change -0.073. It could sit between -0.215 and +0.070.

skills/skill-creator Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.909. Change -0.091. It could sit between -0.269 and +0.087.

skills/skill-creation-walkthrough Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.891. Change -0.109. It could sit between -0.218 and -0.000.

skills/skill-creation-walkthrough Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.855. Change -0.145. It could sit between -0.336 and +0.045.

skills/skill-creation-walkthrough Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.855. Change -0.145. It could sit between -0.336 and +0.045.

skills/skill-creator Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

Scored worse

No skill: 1.000. With the skill: 0.782. Change -0.218. It could sit between -0.436 and -0.000.

skills/skill-creator Claude Fable 5.1 Wider run, three ways, Deterministic

Scored worse

No skill: 1.000. With the skill: 0.564. Change -0.436. It could sit between -0.669 and -0.204.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.
SkillCohortWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loaded Wider run, three waysbaseline76%0 by definition02026-09-05
One-line instruction instead Wider run, three waysbaseline72%0 by definition02026-09-05
skills/skill-creation-walkthrough from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this testrampstackco/claude-skills*76% -> 100%+0.32 [+0.25, +0.38]Stable8,1902026-09-05
skills/skill-creator from anthropics/skills 3b3fad96af16 Wider run, three waysanthropics/skills82% -> 100%+0.30 [+0.25, +0.34]Stable10,8722026-09-05
distribution/claude-plugin/skills/skill-builder from yusufkaraaslan/Skill_Seekers f3972efa33fa Wider run, three waysyusufkaraaslan/Skill_Seekers70% -> 99%+0.26 [+0.25, +0.27]Stable1,2032026-09-05
skills/skill-creator from openclaw/openclaw 61b08f0ebb7a Wider run, three waysopenclaw/openclaw75% -> 99%+0.22 [+0.19, +0.25]Mixed, see page5872026-09-05
No skill loaded First run, with and without the skillbaseline74%0 by definition02026-08-30
One-line instruction instead First run, with and without the skillbaselineinstruction arm: not measured0 by definition02026-08-30
skills/skill-creation-walkthrough from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this testrampstackco/claude-skills*74% -> 100%+0.26 [+0.13, +0.38]Stable on 2 of 38,3902026-08-30
skills/skill-creator from anthropics/skills 3b3fad96af16 First run, with and without the skillanthropics/skills74% -> 100%+0.26 [+0.13, +0.38]Stable on 2 of 311,5142026-08-30
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.