Claude skills for Product Management, tested
Skills for product management work, measured on spec writing tasks. Each was run on the same task pairs with the skill loaded and without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.
Measured on 3 models (Claude, GPT & Gemini): Claude Haiku 4.5, GPT-5 mini, Gemini 3.1 Flash Lite.
The model is given a feature request and asked for a product spec meeting a stated contract: required sections and testable acceptance criteria. Scored two ways: against the contract, and by a grader comparing the with and without outputs blind. Both readings appear.
What helped, and on how many models
First
skills/pm-spec-writing rampstackco/claude-skills*
Spec writing · Paired blind comparison, win rate with ties counting half
Helped clearly on Claude Haiku 4.5, GPT-5 mini and Gemini 3.1 Flash Lite. Measured against no skill loaded.
A judge preferred its output 25 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.
- Claude Haiku 4.5 · helped
- GPT-5 mini · helped
- Gemini 3.1 Flash Lite · helped
Cost per task, with the skill loaded: about $0.00657.Adds about 6,800 tokens of context to every request.
Second
skills/product-capability affaan-m/ECC
Spec writing · Paired blind comparison, win rate with ties counting half
Helped clearly on Claude Haiku 4.5 and Gemini 3.1 Flash Lite; unclear on GPT-5 mini. Measured against no skill loaded.
A judge preferred its output 27 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.
- Claude Haiku 4.5 · helped
- GPT-5 mini · unclear
- Gemini 3.1 Flash Lite · helped
Cost per task, with the skill loaded: about $0.00305.Adds about 1,100 tokens of context to every request.
Tied for third · shared with 1 other, and the tie is broken by nothing
skills/product-capability affaan-m/ECC
Spec writing · Deterministic, scored against a committed answer key
No reliable difference on any model tested. Measured against no skill loaded.
Found 95% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.
Via API
- Claude Haiku 4.5 · unclear
- GPT-5 mini · unclear
- Gemini 3.1 Flash Lite · unclear
In Claude Code
- Claude Fable 5, Skill cohort (one pass) · unclear
- Claude Opus 5, Skill cohort (one pass) · unclear
- Claude Fable 5, Panel v2 (repeat sampling) · unclear
- Claude Haiku 4.5, Panel v2 (repeat sampling) · unclear
- Claude Opus 5, Panel v2 (repeat sampling) · unclear
- Claude Sonnet 5, Panel v2 (repeat sampling) · unclear
- Claude Fable 5.1, Fable 5.1, one pass · unclear
Cost per task, with the skill loaded: about $0.00305.Adds about 1,100 tokens of context to every request.
Tied for third · shared with 1 other, and the tie is broken by nothing
skills/pm-spec-writing rampstackco/claude-skills*
Spec writing · Deterministic, scored against a committed answer key
No reliable difference on any model tested. Measured against no skill loaded.
Found 94% of the required sections and acceptance criteria, against 98% with no skill loaded. Instruction arm: not measured.
Via API
- Claude Haiku 4.5 · unclear
- GPT-5 mini · unclear
- Gemini 3.1 Flash Lite · unclear
In Claude Code
- Claude Fable 5, Skill cohort (one pass) · unclear
- Claude Opus 5, Skill cohort (one pass) · unclear
- Claude Fable 5, Panel v2 (repeat sampling) · unclear
- Claude Haiku 4.5, Panel v2 (repeat sampling) · unclear
- Claude Opus 5, Panel v2 (repeat sampling) · unclear
- Claude Sonnet 5, Panel v2 (repeat sampling) · unclear
- Claude Fable 5.1, Fable 5.1, one pass · unclear
Cost per task, with the skill loaded: about $0.00657.Adds about 6,800 tokens of context to every request.
The numbers
How to read these tables
How to read these tables
- Without -> with
- The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
- Effect (range)
- With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
- Orbit
- Stable: delta at or above the pass threshold, and the interval excludes zero
- Past the horizon: delta at or below the failure floor, and the interval excludes zero
- In free drift: the interval spans zero, or the delta sits between the floor and the threshold
- Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound
A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.
- Cohort
- The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
- Commit
- The exact version tested. Later commits are not covered by this verdict.
- Injected context tokens
- How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
- Tested
- The date. Verdicts age as models change.
Spec writing
| Skill | Cohort | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|
| No skill loaded Paired blind comparison, win rate with ties counting half | baseline | 13% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead Paired blind comparison, win rate with ties counting half | baseline | instruction arm: not applicable to a blind preference | 0 by definition | 0 | 2026-08-30 | |
skills/product-capability from affaan-m/ECC d8409a4b0813 Paired blind comparison, win rate with ties counting half Best result in this test | affaan-m/ECC | 10% -> 90% baseline shown as the complement | +0.80 [+0.40, +1.00] | Stable on 2 of 3 | 1,082 | 2026-08-30 |
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Paired blind comparison, win rate with ties counting half | rampstackco/claude-skills* | 17% -> 83% baseline shown as the complement | +0.67 [+0.60, +0.80] | Stable | 6,845 | 2026-08-30 |
| No skill loaded Deterministic, scored against a committed answer key | baseline | 98% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead Deterministic, scored against a committed answer key | baseline | instruction arm: not measured | 0 by definition | 0 | 2026-08-30 | |
skills/product-capability from affaan-m/ECC d8409a4b0813 Deterministic, scored against a committed answer key Best result in this test | affaan-m/ECC | 98% -> 95% | -0.02 [-0.04, +0.00] | In free drift | 1,082 | 2026-08-30 |
skills/pm-spec-writing from rampstackco/claude-skills 047924252254 Deterministic, scored against a committed answer key | rampstackco/claude-skills* | 98% -> 94% | -0.04 [-0.12, +0.00] | In free drift | 6,845 | 2026-08-30 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | ||||||
Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.