Claude skills for Product Management, tested

Skills for product management work, measured on spec writing tasks. Each was the same task, run twice, once with the skill loaded and once without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.

Measured on 3 models (Claude, GPT & Gemini): Claude Haiku 4.5, GPT-5 mini, Gemini 3.1 Flash Lite.

The model is given a feature request and asked for a product spec meeting a stated contract: required sections and testable acceptance criteria. Scored two ways: against the contract, and by a grader comparing the with and without outputs blind. Both readings appear.

What helped, and on how many models

  • First

    skills/pm-spec-writing rampstackco/claude-skills*

    Spec writing · Blind comparison of the same task, run twice, with ties counting half

    Helped clearly on Claude Haiku 4.5, Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 25 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    Cost per task, with the skill loaded: about $0.00657. Adds about 6,800 tokens of context to every request.

  • Second

    skills/product-capability affaan-m/ECC

    Spec writing · Blind comparison of the same task, run twice, with ties counting half

    Helped clearly on Claude Haiku 4.5 and Gemini 3.1 Flash Lite; unclear on GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 27 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · unclear

    Cost per task, with the skill loaded: about $0.00305. Adds about 1,100 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    skills/product-capability affaan-m/ECC

    Spec writing · Deterministic, scored against a committed answer key

    No reliable difference on any model tested. Measured against the same task with a one-line instruction.

    Found 98% of the required sections and acceptance criteria, against 98% with no skill loaded. Against 97% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5, First run, single attempt · unclear
    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · unclear
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.0022. Adds about 1,000 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    skills/spec-driven-development addyosmani/agent-skills

    Spec writing · Deterministic, scored against a committed answer key

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction.

    Found 98% of the required sections and acceptance criteria, against 98% with no skill loaded. Against 98% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00257. Adds about 2,800 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    pm-execution/skills/create-prd phuryn/pm-skills

    Spec writing · Deterministic, scored against a committed answer key

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction.

    Found 98% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 98% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00197. Adds about 900 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    engineering/skills/spec-driven-workflow alirezarezvani/claude-skills

    Spec writing · Deterministic, scored against a committed answer key

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction.

    Found 97% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 97% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00554. Adds about 15,000 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    skills/product-capability affaan-m/ECC

    Spec writing · Deterministic, scored against a committed answer key

    No reliable difference on any model tested. Measured against no skill loaded.

    Found 95% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · unclear
    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5, First run, single attempt · unclear
    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · unclear
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00305. Adds about 1,100 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    skills/prd github/awesome-copilot

    Spec writing · Deterministic, scored against a committed answer key

    Could not be measured on Gemini 3.1 Flash Lite. Measured against the same task with a one-line instruction.

    Found 95% of the required sections and acceptance criteria, against 97% with no skill loaded. Against 98% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · not measured
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00198. Adds about 1,100 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    skills/pm-spec-writing rampstackco/claude-skills*

    Spec writing · Deterministic, scored against a committed answer key

    No reliable difference on any model tested. Measured against the same task with a one-line instruction.

    Found 95% of the required sections and acceptance criteria, against 99% with no skill loaded. Against 98% with a one-line instruction instead.

    Via API

    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5, First run, single attempt · unclear
    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · unclear
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00357. Adds about 6,700 tokens of context to every request.

  • Tied for third · shared with 7 others, and the tie is broken by nothing

    skills/pm-spec-writing rampstackco/claude-skills*

    Spec writing · Deterministic, scored against a committed answer key

    No reliable difference on any model tested. Measured against no skill loaded.

    Found 94% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.

    Via API

    • Claude Haiku 4.5 · unclear
    • Gemini 3.1 Flash Lite · unclear
    • GPT-5 mini · unclear

    In Claude Code

    • Claude Fable 5, First run, single attempt · unclear
    • Claude Opus 5, First run, single attempt · unclear
    • Claude Fable 5, Second run, every task twice · unclear
    • Claude Haiku 4.5, Second run, every task twice · unclear
    • Claude Opus 5, Second run, every task twice · unclear
    • Claude Sonnet 5, Second run, every task twice · unclear
    • Claude Fable 5.1, Fable 5.1, single attempt · unclear
    • Claude Fable 5.1, Wider run, three ways · unclear
    • Claude Haiku 4.5, Wider run, three ways · unclear
    • Claude Opus 5, Wider run, three ways · unclear
    • Claude Sonnet 5, Wider run, three ways · unclear

    Cost per task, with the skill loaded: about $0.00657. Adds about 6,800 tokens of context to every request.

The numbers

How to read these tables

How to read these tables

Without -> with
The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
Effect (range)
With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
Orbit
  • Stable: delta at or above the pass threshold, and the interval excludes zero
  • Past the horizon: delta at or below the failure floor, and the interval excludes zero
  • In free drift: the interval spans zero, or the delta sits between the floor and the threshold
  • Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound

A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.

Cohort
The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
Commit
The exact version tested. Later commits are not covered by this verdict.
Injected context tokens
How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
Tested
The date. Verdicts age as models change.

Spec writing

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

engineering/skills/spec-driven-workflow Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

pm-execution/skills/create-prd Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

skills/prd Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

skills/product-capability Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.000. With the skill: 1.000. Change +1.000. It could sit between +1.000 and +1.000.

skills/product-capability Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.000. With the skill: 1.000. Change +1.000. It could sit between +1.000 and +1.000.

skills/spec-driven-development Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Could not measure

No skill: 0.991. With the skill: 1.000. Instruction only: 1.000. Change +0.009. It could sit between -0.009 and +0.027.

skills/pm-spec-writing Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Instruction only: 1.000. Change +0.000. It could sit between -0.027 and +0.027.

skills/pm-spec-writing Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Instruction only: 1.000. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability Gemini 3.1 Flash Lite First run, with and without the skill, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Change +0.000. It could sit between -0.027 and +0.027.

skills/pm-spec-writing Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.991. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.973. Instruction only: 0.945. Change +0.000. It could sit between -0.038 and +0.038.

skills/product-capability Claude Haiku 4.5 First run, with and without the skill, Deterministic

No measured effect

No skill: 1.000. With the skill: 0.964. Change -0.036. It could sit between -0.065 and -0.007.

pm-execution/skills/create-prd GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.955. Instruction only: 0.955. Change +0.009. It could sit between -0.023 and +0.041.

skills/spec-driven-development GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 0.955. Instruction only: 0.955. Change -0.009. It could sit between -0.051 and +0.032.

engineering/skills/spec-driven-workflow GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.936. Instruction only: 0.945. Change -0.018. It could sit between -0.054 and +0.017.

skills/pm-spec-writing Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.100. With the skill: 0.900. Change +0.800. It could sit between +0.408 and +1.192.

skills/pm-spec-writing GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.991. With the skill: 0.900. Instruction only: 0.964. Change -0.091. It could sit between -0.156 and -0.026.

skills/product-capability GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.900. Change -0.036. It could sit between -0.175 and +0.102.

skills/prd GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.891. Instruction only: 0.955. Change -0.064. It could sit between -0.234 and +0.107.

skills/pm-spec-writing GPT-5 mini First run, with and without the skill, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.836. Change -0.118. It could sit between -0.254 and +0.017.

skills/pm-spec-writing Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/pm-spec-writing GPT-5 mini First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/product-capability GPT-5 mini First run, with and without the skill, Blind comparison of the same task

No measured effect

No skill: 0.300. With the skill: 0.700. Change +0.400. It could sit between -0.199 and +0.999.

In Claude Code tested in Claude Code

skills/pm-spec-writing Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 1.000. Change +0.036. It could sit between +0.007 and +0.065.

pm-execution/skills/create-prd Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.991. Change +0.036. It could sit between +0.007 and +0.065.

skills/prd Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.991. Change +0.018. It could sit between -0.006 and +0.042.

skills/pm-spec-writing Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.982. With the skill: 0.991. Change +0.009. It could sit between -0.023 and +0.041.

pm-execution/skills/create-prd Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 0.982. Change +0.018. It could sit between -0.017 and +0.054.

skills/pm-spec-writing Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.996. With the skill: 0.977. Change -0.018. It could sit between -0.038 and +0.002.

skills/prd Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.973. Change +0.018. It could sit between -0.026 and +0.063.

skills/product-capability Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.964. With the skill: 0.973. Change +0.009. It could sit between -0.023 and +0.041.

skills/product-capability Claude Haiku 4.5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.977. With the skill: 0.968. Change -0.009. It could sit between -0.035 and +0.017.

skills/product-capability Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.964. Change -0.009. It could sit between -0.041 and +0.023.

skills/pm-spec-writing Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.964. Change -0.009. It could sit between -0.051 and +0.032.

engineering/skills/spec-driven-workflow Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.955. Change -0.018. It could sit between -0.076 and +0.040.

skills/product-capability Claude Fable 5.1 Fable 5.1, single attempt, Deterministic

No measured effect

No skill: 0.973. With the skill: 0.955. Change -0.018. It could sit between -0.054 and +0.017.

pm-execution/skills/create-prd Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.955. Change +0.018. It could sit between -0.026 and +0.063.

skills/pm-spec-writing Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.950. With the skill: 0.950. Change +0.000. It could sit between -0.040 and +0.040.

skills/pm-spec-writing Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.946. Change +0.027. It could sit between +0.004 and +0.051.

skills/prd Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.945. Change -0.009. It could sit between -0.051 and +0.032.

skills/spec-driven-development Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.945. Change +0.000. It could sit between -0.027 and +0.027.

pm-execution/skills/create-prd Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.945. Change +0.036. It could sit between +0.007 and +0.065.

skills/pm-spec-writing Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.945. Change +0.009. It could sit between -0.032 and +0.051.

skills/product-capability Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.945. Change +0.000. It could sit between -0.027 and +0.027.

skills/product-capability Claude Sonnet 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.946. With the skill: 0.941. Change -0.004. It could sit between -0.029 and +0.020.

skills/pm-spec-writing Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.936. Change +0.027. It could sit between -0.019 and +0.074.

skills/pm-spec-writing Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.927. With the skill: 0.936. Change +0.009. It could sit between -0.032 and +0.051.

skills/product-capability Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.927. With the skill: 0.936. Change +0.009. It could sit between -0.023 and +0.041.

skills/product-capability Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.936. Change +0.018. It could sit between -0.017 and +0.054.

skills/spec-driven-development Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.982. With the skill: 0.936. Change -0.045. It could sit between -0.075 and -0.016.

skills/pm-spec-writing Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.946. With the skill: 0.927. Change -0.018. It could sit between -0.045 and +0.009.

skills/product-capability Claude Fable 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.968. With the skill: 0.927. Change -0.041. It could sit between -0.066 and -0.016.

skills/product-capability Claude Opus 5 Second run, every task twice, Deterministic

No measured effect

No skill: 0.932. With the skill: 0.923. Change -0.009. It could sit between -0.035 and +0.017.

skills/product-capability Claude Opus 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.918. Change +0.000. It could sit between -0.027 and +0.027.

skills/pm-spec-writing Claude Fable 5 First run, single attempt, Deterministic

No measured effect

No skill: 0.936. With the skill: 0.918. Change -0.018. It could sit between -0.054 and +0.017.

skills/prd Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.909. Change +0.000. It could sit between +0.000 and +0.000.

skills/spec-driven-development Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.909. Change +0.000. It could sit between +0.000 and +0.000.

skills/spec-driven-development Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.909. Change -0.036. It could sit between -0.065 and -0.007.

engineering/skills/spec-driven-workflow Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.909. With the skill: 0.900. Change -0.009. It could sit between -0.041 and +0.023.

engineering/skills/spec-driven-workflow Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.945. With the skill: 0.900. Change -0.045. It could sit between -0.085 and -0.006.

engineering/skills/spec-driven-workflow Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.955. With the skill: 0.818. Change -0.136. It could sit between -0.166 and -0.107.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.
SkillCohortWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loaded Blind comparison of the same task, run twice, with ties counting half, First run, with and without the skillbaseline13%0 by definition02026-08-30
One-line instruction instead Blind comparison of the same task, run twice, with ties counting half, First run, with and without the skillbaselineinstruction arm: not applicable to a blind preference0 by definition02026-08-30
skills/product-capability from affaan-m/ECC d8409a4b0813 Blind comparison of the same task, run twice, with ties counting half First run, with and without the skill Best result in this testaffaan-m/ECC10% -> 90% baseline shown as the complement+0.80 [+0.40, +1.00]Stable on 2 of 31,0822026-08-30
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Blind comparison of the same task, run twice, with ties counting half First run, with and without the skillrampstackco/claude-skills*17% -> 83% baseline shown as the complement+0.67 [+0.60, +0.80]Stable6,8452026-08-30
No skill loaded Deterministic, scored against a committed answer key, Wider run, three waysbaseline98%0 by definition02026-09-05
One-line instruction instead Deterministic, scored against a committed answer key, Wider run, three waysbaseline98%0 by definition02026-09-05
skills/product-capability from affaan-m/ECC d8409a4b0813 Deterministic, scored against a committed answer key Wider run, three ways Best result in this testaffaan-m/ECC98% -> 98%+0.01 [-0.01, +0.03]In free drift1,0472026-09-05
skills/spec-driven-development from addyosmani/agent-skills d2c37ef6225d Deterministic, scored against a committed answer key Wider run, three waysaddyosmani/agent-skills98% -> 98%+0.00 [+0.00, +0.00]Mixed, see page2,7772026-09-05
pm-execution/skills/create-prd from phuryn/pm-skills 18468a95b427 Deterministic, scored against a committed answer key Wider run, three waysphuryn/pm-skills97% -> 98%+0.00 [+0.00, +0.00]Mixed, see page9412026-09-05
engineering/skills/spec-driven-workflow from alirezarezvani/claude-skills 19392f7a0826 Deterministic, scored against a committed answer key Wider run, three waysalirezarezvani/claude-skills97% -> 97%-0.00 [-0.01, +0.00]Mixed, see page15,0432026-09-05
skills/prd from github/awesome-copilot c956566a35c3 Deterministic, scored against a committed answer key Wider run, three waysgithub/awesome-copilot97% -> 95%-0.03 [-0.06, +0.00]Mixed, see page1,1472026-09-05
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Deterministic, scored against a committed answer key Wider run, three waysrampstackco/claude-skills*99% -> 95%-0.04 [-0.06, -0.01]In free drift6,6532026-09-05
No skill loaded Deterministic, scored against a committed answer key, First run, with and without the skillbaseline98%0 by definition02026-08-30
One-line instruction instead Deterministic, scored against a committed answer key, First run, with and without the skillbaselineinstruction arm: not measured0 by definition02026-08-30
skills/product-capability from affaan-m/ECC d8409a4b0813 Deterministic, scored against a committed answer key First run, with and without the skill Best result in this testaffaan-m/ECC98% -> 95%-0.02 [-0.04, +0.00]In free drift1,0822026-08-30
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Deterministic, scored against a committed answer key First run, with and without the skillrampstackco/claude-skills*98% -> 94%-0.04 [-0.12, +0.00]In free drift6,8452026-08-30
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.