Does PRD help? Tested on spec writing tasks
PRD, from github/awesome-copilot. Its spec writing tasks were run with the skill loaded and without it. Each time it is the same task, run twice.
What we tested
Whether loading this skill helps a model write a product spec.
What counts as helping
The score with the skill has to beat the score without it by at least 10 percentage points, on the same 10 briefs. A blind reader has to prefer the version written with the skill at least 60 times in 100, across the same 10 briefs.
How we scored it
We checked each spec against a fixed list of things a spec needs, and we also had a blind reader compare the two versions. The two are reported separately and never added together.
All 6 models tested already passed this without the change, and still passed with it.
There is no chart here because there is no shape to draw: every value sits at one end of the scale, the models agree, and no range of likely values is wide enough to see. The per-model numbers are in the tables below.
Show per-task detail
Every item, without and with, per model
Via API tested via API
- GPT-5 mini Deterministic
- Gemini 3.1 Flash Lite Deterministic
deterministic pass rate, 0 to 1
Hover or focus an item to read its task and every model’s two values.
Task-paired: these item means are what this cell's interval is built from.
The item means behind this plot
| Item | GPT-5 mini Deterministic | Gemini 3.1 Flash Lite Deterministic | ||
|---|---|---|---|---|
| Task | without | with | without | with |
| sw-01 | 0.909 | 0.909 | 1.000 | 1.000 |
| sw-03 | 0.909 | 0.182 | 1.000 | 1.000 |
| sw-05 | 0.909 | 0.909 | 1.000 | 1.000 |
| sw-07 | 0.909 | 1.000 | 1.000 | 1.000 |
| sw-08 | 0.909 | 1.000 | 1.000 | 1.000 |
| sw-02 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-04 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-06 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-09 | 1.000 | 1.000 | 1.000 | 1.000 |
| sw-10 | 1.000 | 0.909 | 1.000 | 1.000 |
Without to with, per model
In Claude Code tested in Claude Code
Claude Opus 5 Expansion cohort (three arms)
No measured effect
Claude Fable 5.1 Expansion cohort (three arms)
No measured effect
Claude Sonnet 5 Expansion cohort (three arms)
No measured effect
Claude Haiku 4.5 Expansion cohort (three arms)
No measured effect
deterministic pass rate, 0 to 1
The Claude Code runs committed their cells as arm means and intervals rather than as per-item values, so this panel is drawn without to with rather than item by item.
What exactly was tested, and how it was scored
- Repository
- github/awesome-copilot
- Path
- skills/prd
- Commit
c956566a35c3c2e635f019e7a1bfa59d9497e8b1- Content hash
fc1caf7b8c27e187d50b76ae52b0972bbf95b7bbc16e39a4d138a132bd76aea8- Date tested
- not yet tested
The pin is the whole of this skill's identity here. It resolves at https://github.com/github/awesome-copilot/tree/c956566a35c3c2e635f019e7a1bfa59d9497e8b1/skills/prd, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.
S19-github-prd
Loading skills/prd from github/awesome-copilot improves outputs on spec writing tasks.
- Pass criterion
- With-arm mean structure score exceeds the without-arm by at least 0.10, and the with-arm win rate is at or above 0.60, on the same 10 briefs.
- Scale
- unit. Same scorer, same instrument and same scale as S06-product-capability. Added 2026-08-31 by the source expansion under R23; adding a row to an instrument does not touch the instrument.
- Task pairs planned per model
- 10
- Notes
- Shares the spec-writing task set with every other claim in this class, so all of its rows are measured on identical items against one answer key. The slot this claim scores was pinned by the 2026-08-31 source expansion under R23 and NO MODEL HAS BEEN RUN AGAINST IT. The claim exists so the class holds one claim per pinned slot, which is what makes the row comparable the day it is run.
Injected context tokens
| Arm | Context characters | Injected context tokens |
|---|---|---|
| without | 0 | 0 |
| with | 4,328 | 1,016 |
The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.
Verdict per model
| Model | Reading | Orbit | Without -> with | Effect | 95 percent interval | Task pairs | Model version returned |
|---|---|---|---|---|---|---|---|
| Gemini 3.1 Flash Lite | Deterministic, scored against a committed answer key | Unobservable (b) both arms at a bound, the task set did not separate them | 1.0000 -> 1.0000 | +0.0000 | [0.0000, 0.0000] | 10 pairs, 30 of 30 records | gemini-3.1-flash-lite |
| GPT-5 mini | Deterministic, scored against a committed answer key | In free drift (10 of 10 pairs) Nothing left to measure. The model scored 95% without the skill. | 0.9545 -> 0.8909 | -0.0636 | [-0.2116, 0.0844] | 10 pairs, 30 of 32 records | gpt-5-mini-2025-08-07, gpt-5-mini |
- Unobservable: the cell cannot be classified: an unbounded metric, or both arms at a bound of the scale.
- In free drift: no separation the design can resolve. Not evidence of no effect.
Cost
Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.
Where the with arm did not win
Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key
Drew on 10 of 10 pairs: sw-01, sw-02, sw-03, sw-04, sw-05, sw-06, sw-07, sw-08, sw-09, sw-10.
GPT-5 mini, Deterministic, scored against a committed answer key
Lost on 2 of 10 pairs: sw-03 (-0.7273), sw-10 (-0.0909).
Drew on 6 of 10 pairs: sw-01, sw-02, sw-04, sw-05, sw-06, sw-09.
In Claude Code
These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument: no figure here is averaged with one tested via API above. How the two were compared.
- Claude Fable 5.1 · unclear in the three-arm cohort
- Claude Haiku 4.5 · unclear in the three-arm cohort
- Claude Opus 5 · unclear in the three-arm cohort
- Claude Sonnet 5 · unclear in the three-arm cohort
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 0.9545 | 0.9455 | -0.009 | -0.051 to 0.032 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
| Claude Haiku 4.5 | 0.9727 | 0.9909 | +0.018 | -0.006 to 0.042 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
| Claude Opus 5 | 0.9091 | 0.9091 | +0.000 | 0.000 to 0.000 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
| Claude Sonnet 5 | 0.9545 | 0.9727 | +0.018 | -0.026 to 0.063 | 10 | In free drift | not measured: this run records the dose per arm, in the doses list |
Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.