Does PM spec writing help? Tested on spec writing tasks
PM spec writing, from rampstackco/claude-skills. Its spec writing tasks were run with the skill loaded and without it. Each time it is the same task, run twice.
What we tested
Whether loading this skill helps a model write a product spec.
What counts as helping
The score with the skill has to beat the score without it by at least 10 percentage points, on the same 10 briefs. A blind reader has to prefer the version written with the skill at least 60 times in 100, across the same 10 briefs.
How we scored it
We checked each spec against a fixed list of things a spec needs, and we also had a blind reader compare the two versions. The two are reported separately and never added together.
Without to with, per model
Via API tested via API
Gemini 3.1 Flash Lite Deterministic, First run, with and without the skill
No measured effect
Gemini 3.1 Flash Lite Deterministic, Wider run, three ways
No measured effect
Claude Haiku 4.5 Deterministic, First run, with and without the skill
No measured effect
Claude Haiku 4.5 Blind comparison of the same task, First run, with and without the skill
Holds
GPT-5 mini Deterministic, Wider run, three ways
No measured effect
GPT-5 mini Deterministic, First run, with and without the skill
No measured effect
Gemini 3.1 Flash Lite Blind comparison of the same task, First run, with and without the skill
Holds
GPT-5 mini Blind comparison of the same task, First run, with and without the skill
Holds
deterministic pass rate, 0 to 1
In Claude Code tested in Claude Code
Claude Fable 5.1 Wider run, three ways
No measured effect
Claude Fable 5.1 Fable 5.1, single attempt
No measured effect
Claude Fable 5 Second run, every task twice
No measured effect
Claude Fable 5 First run, single attempt
No measured effect
Claude Haiku 4.5 Second run, every task twice
No measured effect
Claude Haiku 4.5 Wider run, three ways
No measured effect
Claude Sonnet 5 Second run, every task twice
No measured effect
Claude Sonnet 5 Wider run, three ways
No measured effect
Claude Opus 5 Second run, every task twice
No measured effect
Claude Opus 5 First run, single attempt
No measured effect
Claude Opus 5 Wider run, three ways
No measured effect
deterministic pass rate, 0 to 1
Show per-task detail
Every item, without and with, per model
Via API tested via API
- Claude Haiku 4.5 Blind comparison of the same task
- Gemini 3.1 Flash Lite Blind comparison of the same task
- GPT-5 mini Blind comparison of the same task
- GPT-5 mini Deterministic
- GPT-5 mini Deterministic
- Claude Haiku 4.5 Deterministic
- Gemini 3.1 Flash Lite Deterministic
- Gemini 3.1 Flash Lite Deterministic
deterministic pass rate, 0 to 1
Hover or focus an item to read its task and every model’s two values.
Task-paired: these item means are what this cell's interval is built from.
The item means behind this plot
| Item | Claude Haiku 4.5 Blind comparison of the same task | Gemini 3.1 Flash Lite Blind comparison of the same task | GPT-5 mini Blind comparison of the same task | GPT-5 mini Deterministic | GPT-5 mini Deterministic | Claude Haiku 4.5 Deterministic | Gemini 3.1 Flash Lite Deterministic | Gemini 3.1 Flash Lite Deterministic | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Task | without | with | without | with | without | with | without | with | without | with | without | with | without | with | without | with |
| sw-04 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 | 0.909 | 0.909 | 0.909 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-09 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 | 0.909 | 0.909 | 0.909 | 1.000 | 1.000 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-07 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 | 1.000 | 0.909 | 1.000 | 1.000 | 1.000 |
| sw-08 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 | 0.909 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-03 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 | 1.000 | 1.000 | 0.909 | 1.000 | 0.909 |
| sw-06 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-01 | 0.000 | 1.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.909 | 0.909 | 0.909 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-05 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.000 | 1.000 | 0.909 | 0.909 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-10 | 0.000 | 1.000 | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 | 0.909 | 1.000 | 0.636 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-02 | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 | 0.000 | 0.909 | 0.182 | 1.000 | 0.909 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
Without to with, per model
In Claude Code tested in Claude Code
Claude Opus 5 Expansion cohort (three arms)
No measured effect
Claude Opus 5 Panel v2 (repeat sampling)
No measured effect
Claude Opus 5 Skill cohort (one pass)
No measured effect
Claude Fable 5 Skill cohort (one pass)
No measured effect
Claude Sonnet 5 Expansion cohort (three arms)
No measured effect
Claude Fable 5 Panel v2 (repeat sampling)
No measured effect
Claude Sonnet 5 Panel v2 (repeat sampling)
No measured effect
Claude Fable 5.1 Expansion cohort (three arms)
No measured effect
Claude Haiku 4.5 Expansion cohort (three arms)
No measured effect
Claude Fable 5.1 Fable 5.1, one pass
No measured effect
Claude Haiku 4.5 Panel v2 (repeat sampling)
No measured effect
deterministic pass rate, 0 to 1
The Claude Code runs committed their cells as arm means and intervals rather than as per-item values, so this panel is drawn without to with rather than item by item.
What exactly was tested, and how it was scored
- Repository
- rampstackco/claude-skills
- Path
- skills/pm-spec-writing
- Commit
0479242522549dfdb389bb9b7807ad4d6016ffb7- Content hash
79637738ca6291618e0ec2987efdb6759d5bfda223101bf795001ca6d4b28552- Date tested
- 2026-08-30
The pin is the whole of this skill's identity here. It resolves at https://github.com/rampstackco/claude-skills/tree/0479242522549dfdb389bb9b7807ad4d6016ffb7/skills/pm-spec-writing, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.
S05-pm-spec-writing
Loading skills/pm-spec-writing from rampstackco/claude-skills improves outputs on spec writing tasks.
- Pass criterion
- With-arm mean structure score exceeds the without-arm by at least 0.10, and the with-arm win rate is at or above 0.60, on the same 10 briefs.
- Scale
- unit. The deterministic half is a fraction of structure checks passed, a [0,1] fraction. The paired half is a win rate, also bounded at 0 and 1. Both are the unit scale and the 0.2 threshold applies to each.
- Task pairs planned per model
- 10
- Notes
- The only class scored twice. The structure check and the blind comparison are reported as two separate cells and are never averaged into one number: a spec can have every heading and be useless, and a scorer that mixed the two would let one hide the other.
Injected context tokens
| Arm | Context characters | Injected context tokens |
|---|---|---|
| without | 0 | 0 |
| with | 26,805 | 6,292 |
The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.
Verdict per model
| Model | Reading | Orbit | Without -> with | Effect | 95 percent interval | Task pairs | Model version returned |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | Deterministic, scored against a committed answer key | In free drift Nothing left to measure. The model scored 99% without the skill. | 0.9909 -> 0.9909 | +0.0000 | [-0.0266, 0.0266] | 10 pairs, 20 of 20 records | claude-haiku-4-5-20251001 |
| Claude Haiku 4.5 | Blind comparison of the same task, run twice, with ties counting half | Stable | 0.1000 -> 0.9000 baseline shown as the complement | win rate 0.9000, +0.8000 on the arm difference | [0.4080, 1.1920] | 10 pairs, 20 of 20 records | claude-haiku-4-5-20251001 |
| Gemini 3.1 Flash Lite | Deterministic, scored against a committed answer key | In free drift Nothing left to measure. The model scored 99% without the skill. | 0.9909 -> 0.9909 | +0.0000 | [-0.0266, 0.0266] | 10 pairs, 20 of 20 records | gemini-3.1-flash-lite |
| Gemini 3.1 Flash Lite | Blind comparison of the same task, run twice, with ties counting half | Stable | 0.2000 -> 0.8000 baseline shown as the complement | win rate 0.8000, +0.6000 on the arm difference | [0.0773, 1.1227] | 10 pairs, 20 of 20 records | gemini-3.1-flash-lite |
| GPT-5 mini | Deterministic, scored against a committed answer key | In free drift Nothing left to measure. The model scored 95% without the skill. | 0.9545 -> 0.8364 | -0.1182 | [-0.2538, 0.0174] | 10 pairs, 20 of 20 records | gpt-5-mini-2025-08-07 |
| GPT-5 mini | Blind comparison of the same task, run twice, with ties counting half | Stable | 0.2000 -> 0.8000 baseline shown as the complement | win rate 0.8000, +0.6000 on the arm difference | [0.0773, 1.1227] | 10 pairs, 20 of 20 records | gpt-5-mini-2025-08-07 |
| Gemini 3.1 Flash Lite | Deterministic, scored against a committed answer key | In free drift Nothing left to measure. The model scored 99% without the skill. | 1.0000 -> 0.9909 | -0.0091 | [-0.0269, 0.0087] | 10 pairs, 30 of 30 records | gemini-3.1-flash-lite |
| GPT-5 mini | Deterministic, scored against a committed answer key | In free drift Nothing left to measure. The model scored 99% without the skill. | 0.9636 -> 0.9000 | -0.0636 | [-0.1390, 0.0117] | 10 pairs, 30 of 30 records | gpt-5-mini-2025-08-07 |
- In free drift: no separation the design can resolve. Not evidence of no effect.
- Stable: the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero.
Cost
| Model | Arm | Tasks attempted | Mean input tokens | Mean output tokens | Cost per task | Cost ratio | Grading cost per task pair |
|---|---|---|---|---|---|---|---|
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | without | 10 | 118 | 644 | $0.00334 | ||
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | with | 10 | 7,230 | 1,045 | $0.01245 | 3.73x | $0.00843 |
GPT-5 mini served gpt-5-mini-2025-08-07 | without | 10 | 107 | 1,261 | $0.00255 | ||
GPT-5 mini served gpt-5-mini-2025-08-07 | with | 10 | 6,462 | 1,577 | $0.00477 | 1.87x | $0.01457 |
| Gemini 3.1 Flash Lite | without | 10 | 106 | 513 | $0.00080 | ||
| Gemini 3.1 Flash Lite | with | 10 | 6,843 | 514 | $0.00248 | 3.12x | $0.00570 |
Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.
Where the with arm did not win
Claude Haiku 4.5, Deterministic, scored against a committed answer key
Lost on 1 of 10 pairs: sw-09 (-0.0909).
Drew on 8 of 10 pairs: sw-01, sw-03, sw-04, sw-05, sw-06, sw-07, sw-08, sw-10.
Claude Haiku 4.5, Blind comparison of the same task, run twice, with ties counting half
Lost on 1 of 10 pairs: sw-02 (-1.0000).
Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key
Lost on 1 of 10 pairs: sw-03 (-0.0909).
Drew on 8 of 10 pairs: sw-01, sw-02, sw-04, sw-05, sw-06, sw-08, sw-09, sw-10.
Gemini 3.1 Flash Lite, Blind comparison of the same task, run twice, with ties counting half
Lost on 2 of 10 pairs: sw-01 (-1.0000), sw-10 (-1.0000).
GPT-5 mini, Deterministic, scored against a committed answer key
Lost on 6 of 10 pairs: sw-02 (-0.7273), sw-03 (-0.0909), sw-05 (-0.0909), sw-06 (-0.0909), sw-07 (-0.0909), sw-10 (-0.0909).
Drew on 4 of 10 pairs: sw-01, sw-04, sw-08, sw-09.
GPT-5 mini, Blind comparison of the same task, run twice, with ties counting half
Lost on 2 of 10 pairs: sw-02 (-1.0000), sw-05 (-1.0000).
Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key
Lost on 1 of 10 pairs: sw-03 (-0.0909).
Drew on 9 of 10 pairs: sw-01, sw-02, sw-04, sw-05, sw-06, sw-07, sw-08, sw-09, sw-10.
GPT-5 mini, Deterministic, scored against a committed answer key
Lost on 5 of 10 pairs: sw-10 (-0.3636), sw-02 (-0.0909), sw-03 (-0.0909), sw-06 (-0.0909), sw-07 (-0.0909).
Drew on 4 of 10 pairs: sw-01, sw-04, sw-05, sw-08.
In Claude Code
These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument: no figure here is averaged with one tested via API above. How the two were compared.
- Claude Fable 5 · unclear, and both runs agree
- Claude Opus 5 · unclear, and both runs agree
- Claude Haiku 4.5 · unclear under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · unclear in the version re-test
Skill cohort (one pass)
one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 0.9364 | 0.9182 | -0.018 | -0.054 to 0.017 | 10 | In free drift | 9,996 |
| Claude Opus 5 | 0.9273 | 0.9364 | +0.009 | -0.032 to 0.051 | 10 | In free drift | 9,996 |
| Model | Arm | Records that reasoned | Mean thinking tokens | Most on one record |
|---|---|---|---|---|
| Claude Fable 5 | with skill | 20 of 20 | 43 | 106 |
| Claude Fable 5 | no skill | 20 of 20 | 20 | 37 |
Panel v2 (repeat sampling)
two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 0.9455 | 0.9273 | -0.018 | -0.045 to 0.009 | 10 | In free drift | not measured: 12 of 40 records predate the usage-capture amendment |
| Claude Haiku 4.5 | 0.9955 | 0.9773 | -0.018 | -0.038 to 0.002 | 10 | In free drift | not measured: 28 of 40 records predate the usage-capture amendment |
| Claude Opus 5 | 0.9182 | 0.9455 | +0.027 | 0.004 to 0.051 | 10 | In free drift | not measured: 16 of 40 records predate the usage-capture amendment |
| Claude Sonnet 5 | 0.9500 | 0.9500 | +0.000 | -0.040 to 0.040 | 10 | In free drift | not measured: 20 of 40 records predate the usage-capture amendment |
| Model | Arm | Records that reasoned | Mean thinking tokens | Most on one record |
|---|---|---|---|---|
| Claude Fable 5 | with skill | 20 of 20 | 43 | 106 |
| Claude Fable 5 | no skill | 20 of 20 | 20 | 37 |
Fable 5.1, one pass
one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 0.9818 | 0.9909 | +0.009 | -0.023 to 0.041 | 20 | In free drift | not measured |
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 0.9636 | 1.0000 | +0.036 | 0.007 to 0.065 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
| Claude Haiku 4.5 | 0.9727 | 0.9636 | -0.009 | -0.051 to 0.032 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
| Claude Opus 5 | 0.9091 | 0.9364 | +0.027 | -0.019 to 0.074 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
| Claude Sonnet 5 | 0.9364 | 0.9455 | +0.009 | -0.032 to 0.051 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.