skills/pm-spec-writing from rampstackco/claude-skills

Spec writing tasks, run with the skill loaded and without it, on the same task pairs.

Identity pin

Repository
rampstackco/claude-skills
Path
skills/pm-spec-writing
Commit
0479242522549dfdb389bb9b7807ad4d6016ffb7
Content hash
79637738ca6291618e0ec2987efdb6759d5bfda223101bf795001ca6d4b28552
Date tested
not yet tested

The pin is the whole of this skill's identity here. It resolves at https://github.com/rampstackco/claude-skills/tree/0479242522549dfdb389bb9b7807ad4d6016ffb7/skills/pm-spec-writing, and the content hash is a sha256 over exactly the text the with arm was given. Nothing else about the skill appears on this site.

What was claimed

  • Loading skills/pm-spec-writing from rampstackco/claude-skills improves outputs on spec writing tasks.
    Pass criterion: With-arm mean structure score exceeds the without-arm by at least 0.10, and the with-arm win rate is at or above 0.60, on the same 10 briefs. Scale: unit. The deterministic half is a fraction of structure checks passed, a [0,1] fraction. The paired half is a win rate, also bounded at 0 and 1. Both are the unit scale and the 0.2 threshold applies to each.
    The only class scored twice. The structure check and the blind comparison are reported as two separate cells and are never averaged into one number: a spec can have every heading and be useless, and a scorer that mixed the two would let one hide the other.

Injected context tokens

Injected context tokens, per arm
ArmContext charactersInjected context tokens
without00
with26,8056,292

The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.

Not yet measured

No run records exist for this skill. The instrument is built and committed, the task sets are frozen, and the projected cost of the run is published, but no model has been called. There is therefore no verdict, no effect, no interval and no cost per task on this page, and none is shown.

  • S05-pm-spec-writing: 10 task pairs per model. Scored twice, deterministically and by paired blind comparison, reported as two rows and never averaged..

The committed instrument is at harness/claims/claims-skills.json, harness/claims/tasksets-skills.json and harness/claims/skill-cohort.json. The projection for the run that has not happened is at harness/results/dry-run-projection-skills.md. How a claim gets tested.

0 run records behind this page. Every verdict is computed at build time by the same frozen status_v1 rule that decides every other verdict on this site, and nothing here is written by hand.