Does Product capability help? Tested on spec writing tasks

Product capability, from affaan-m/ECC. Its spec writing tasks were run with the skill loaded and without it. Each time it is the same task, run twice.

What we tested

Whether loading this skill helps a model write a product spec.

What counts as helping

The score with the skill has to beat the score without it by at least 10 percentage points, on the same 10 briefs. A blind reader has to prefer the version written with the skill at least 60 times in 100, across the same 10 briefs.

How we scored it

We checked each spec against a fixed list of things a spec needs, and we also had a blind reader compare the two versions. The two are reported separately and never added together.

Without to with, per model

Via API tested via API

Claude Haiku 4.5 Blind comparison of the same task, First run, with and without the skill

0.000
1.000

Holds

Claude Haiku 4.5 Deterministic, First run, with and without the skill

1.000
0.964

No measured effect

Gemini 3.1 Flash Lite Blind comparison of the same task, First run, with and without the skill

0.000
1.000

Holds

Gemini 3.1 Flash Lite Deterministic, First run, with and without the skill

0.991 without and with

No measured effect

Gemini 3.1 Flash Lite Deterministic, Wider run, three ways

1.000
0.991

No measured effect

GPT-5 mini Deterministic, Wider run, three ways

0.945
0.973

No measured effect

GPT-5 mini Deterministic, First run, with and without the skill

0.936
0.900

No measured effect

GPT-5 mini Blind comparison of the same task, First run, with and without the skill

0.300
0.700

No measured effect

deterministic pass rate, 0 to 1

In Claude Code tested in Claude Code

Claude Fable 5.1 Wider run, three ways

0.964
0.973

No measured effect

Claude Fable 5.1 Fable 5.1, single attempt

0.973
0.955

No measured effect

Claude Fable 5 First run, single attempt

0.927
0.936

No measured effect

Claude Fable 5 Second run, every task twice

0.968
0.927

No measured effect

Claude Haiku 4.5 Second run, every task twice

0.977
0.968

No measured effect

Claude Haiku 4.5 Wider run, three ways

0.973
0.964

No measured effect

Claude Sonnet 5 Wider run, three ways

0.945 without and with

No measured effect

Claude Sonnet 5 Second run, every task twice

0.946
0.941

No measured effect

Claude Opus 5 Wider run, three ways

0.918
0.936

No measured effect

Claude Opus 5 Second run, every task twice

0.932
0.923

No measured effect

Claude Opus 5 First run, single attempt

0.918 without and with

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Grey is the score without. Colour is the score with. The bracket shows how much the difference could move if we ran it again. Rows are ordered by the score with. A model measured at two versions keeps its versions next to each other. The two test methods are reported separately and never averaged. Where the two scores are the same, the number is printed once at the end of the pair. A bracket wider than the axis is drawn to the edge with its cap omitted; the table gives its bounds.
Show per-task detail

Every item, without and with, per model

Via API tested via API

  • Claude Haiku 4.5 Blind comparison of the same task
  • Gemini 3.1 Flash Lite Blind comparison of the same task
  • GPT-5 mini Blind comparison of the same task
  • GPT-5 mini Deterministic
  • GPT-5 mini Deterministic
  • Gemini 3.1 Flash Lite Deterministic
  • Claude Haiku 4.5 Deterministic
  • Gemini 3.1 Flash Lite Deterministic

deterministic pass rate, 0 to 1

Hover or focus an item to read its task and every model’s two values.

Task-paired: these item means are what this cell's interval is built from.

The item means behind this plot
Per-item arm means. Without is the unaided arm, with is the treated arm.
ItemClaude Haiku 4.5 Blind comparison of the same task Gemini 3.1 Flash Lite Blind comparison of the same task GPT-5 mini Blind comparison of the same task GPT-5 mini Deterministic GPT-5 mini Deterministic Gemini 3.1 Flash Lite Deterministic Claude Haiku 4.5 Deterministic Gemini 3.1 Flash Lite Deterministic
Task without with without with without with without with without with without with without with without with
sw-07 0.000 1.000 0.000 1.000 0.000 1.000 0.727 0.909 0.909 0.909 0.909 1.000 1.000 1.000 1.000 1.000
sw-03 0.000 1.000 0.000 1.000 0.000 1.000 0.909 1.000 0.909 0.909 1.000 0.909 1.000 1.000 1.000 0.909
sw-09 0.000 1.000 0.000 1.000 0.000 1.000 0.909 1.000 0.909 0.909 1.000 1.000 1.000 0.909 1.000 1.000
sw-02 0.000 1.000 0.000 1.000 0.000 1.000 0.909 0.273 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
sw-08 0.000 1.000 0.000 1.000 0.000 1.000 1.000 1.000 0.909 1.000 1.000 1.000 1.000 1.000 1.000 1.000
sw-10 0.000 1.000 0.000 1.000 0.000 1.000 1.000 1.000 0.909 1.000 1.000 1.000 1.000 0.909 1.000 1.000
sw-04 0.000 1.000 0.000 1.000 0.000 1.000 1.000 0.909 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
sw-05 0.000 1.000 0.000 1.000 1.000 0.000 1.000 1.000 0.909 1.000 1.000 1.000 1.000 1.000 1.000 1.000
sw-06 0.000 1.000 0.000 1.000 1.000 0.000 0.909 0.909 1.000 1.000 1.000 1.000 1.000 0.909 1.000 1.000
sw-01 0.000 1.000 0.000 1.000 1.000 0.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 0.909 1.000 1.000
A picture of the per-item means behind the per-model numbers. The tables are the source. Each model draws two lines over the same task set: a dashed line through its unaided scores and a solid line through its treated ones. Items are ordered by the mean unaided score across the models that measured them, lowest first. The two instruments are reported separately and never averaged.

Without to with, per model

In Claude Code tested in Claude Code

Claude Opus 5 Expansion cohort (three arms)

No measured effect

Claude Opus 5 Skill cohort (one pass)

No measured effect

Claude Fable 5 Skill cohort (one pass)

No measured effect

Claude Opus 5 Panel v2 (repeat sampling)

No measured effect

Claude Sonnet 5 Expansion cohort (three arms)

No measured effect

Claude Sonnet 5 Panel v2 (repeat sampling)

No measured effect

Claude Fable 5.1 Expansion cohort (three arms)

No measured effect

Claude Fable 5 Panel v2 (repeat sampling)

No measured effect

Claude Fable 5.1 Fable 5.1, one pass

No measured effect

Claude Haiku 4.5 Expansion cohort (three arms)

No measured effect

Claude Haiku 4.5 Panel v2 (repeat sampling)

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Each row runs from the score without to the score with. The thin bar beneath it is the bracket, and it shows how much the difference could move if we ran it again. The two test methods are reported separately and never averaged. A bracket wider than the axis is drawn to the edge with its end cap omitted; the table gives its bounds.

The Claude Code runs committed their cells as arm means and intervals rather than as per-item values, so this panel is drawn without to with rather than item by item.

What exactly was tested, and how it was scored
Repository
affaan-m/ECC
Path
skills/product-capability
Commit
d8409a4b0813771235555e32e3d8046a73988bfa
Content hash
3e0052802969388ce566464adf7ce89302cb99165bd7d759d9b2e22594c7b18c
Date tested
2026-08-30

The pin is the whole of this skill's identity here. It resolves at https://github.com/affaan-m/ECC/tree/d8409a4b0813771235555e32e3d8046a73988bfa/skills/product-capability, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.

S06-product-capability

Loading skills/product-capability from affaan-m/ECC improves outputs on spec writing tasks.

Pass criterion
With-arm mean structure score exceeds the without-arm by at least 0.10, and the with-arm win rate is at or above 0.60, on the same 10 briefs.
Scale
unit. Same two scorers and the same scale as S05.
Task pairs planned per model
10
Notes
Shares the spec-writing task set with S05.

Injected context tokens

Injected context tokens, per arm
ArmContext charactersInjected context tokens
without00
with4,4571,046

The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.

Verdict per model

Effect per model
ModelReadingOrbitWithout -> withEffect95 percent intervalTask pairsModel version returned
Claude Haiku 4.5Deterministic, scored against a committed answer keyIn free drift Nothing left to measure. The model scored 100% without the skill.1.0000 -> 0.9636-0.0364[-0.0655, -0.0073]10 pairs, 20 of 20 recordsclaude-haiku-4-5-20251001
Claude Haiku 4.5Blind comparison of the same task, run twice, with ties counting halfStable0.0000 -> 1.0000 baseline shown as the complementwin rate 1.0000, +1.0000 on the arm difference[1.0000, 1.0000]10 pairs, 20 of 20 recordsclaude-haiku-4-5-20251001
Gemini 3.1 Flash LiteDeterministic, scored against a committed answer keyIn free drift Nothing left to measure. The model scored 99% without the skill.0.9909 -> 0.9909+0.0000[-0.0266, 0.0266]10 pairs, 20 of 20 recordsgemini-3.1-flash-lite
Gemini 3.1 Flash LiteBlind comparison of the same task, run twice, with ties counting halfStable0.0000 -> 1.0000 baseline shown as the complementwin rate 1.0000, +1.0000 on the arm difference[1.0000, 1.0000]10 pairs, 20 of 20 recordsgemini-3.1-flash-lite
GPT-5 miniDeterministic, scored against a committed answer keyIn free drift Nothing left to measure. The model scored 94% without the skill.0.9364 -> 0.9000-0.0364[-0.1749, 0.1022]10 pairs, 20 of 20 recordsgpt-5-mini-2025-08-07
GPT-5 miniBlind comparison of the same task, run twice, with ties counting halfIn free drift unclear. The model scored 30% without it.0.3000 -> 0.7000 baseline shown as the complementwin rate 0.7000, +0.4000 on the arm difference[-0.1988, 0.9988]10 pairs, 20 of 20 recordsgpt-5-mini-2025-08-07
Gemini 3.1 Flash LiteDeterministic, scored against a committed answer keyIn free drift Nothing left to measure. The model scored 99% without the skill.1.0000 -> 0.9909-0.0091[-0.0269, 0.0087]10 pairs, 30 of 30 recordsgemini-3.1-flash-lite
GPT-5 miniDeterministic, scored against a committed answer keyIn free drift Nothing left to measure. The model scored 97% without the skill.0.9455 -> 0.9727+0.0273[0.0001, 0.0545]10 pairs, 30 of 30 recordsgpt-5-mini-2025-08-07
  • In free drift: no separation the design can resolve. Not evidence of no effect.
  • Stable: the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero.

Cost

Cost per task
ModelArmTasks attemptedMean input tokensMean output tokensCost per taskCost ratioGrading cost per task pair
Claude Haiku 4.5 served claude-haiku-4-5-20251001without10118637$0.00330
Claude Haiku 4.5 served claude-haiku-4-5-20251001with101,153743$0.004871.47x$0.00710
GPT-5 mini served gpt-5-mini-2025-08-07without101071,193$0.00241
GPT-5 mini served gpt-5-mini-2025-08-07with101,0261,477$0.003211.33x$0.01391
Gemini 3.1 Flash Litewithout10106513$0.00080
Gemini 3.1 Flash Litewith101,067529$0.001061.33x$0.00582

Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.

Where the with arm did not win

Claude Haiku 4.5, Deterministic, scored against a committed answer key

Lost on 4 of 10 pairs: sw-01 (-0.0909), sw-06 (-0.0909), sw-09 (-0.0909), sw-10 (-0.0909).

Drew on 6 of 10 pairs: sw-02, sw-03, sw-04, sw-05, sw-07, sw-08.

Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key

Lost on 1 of 10 pairs: sw-03 (-0.0909).

Drew on 8 of 10 pairs: sw-01, sw-02, sw-04, sw-05, sw-06, sw-08, sw-09, sw-10.

GPT-5 mini, Deterministic, scored against a committed answer key

Lost on 2 of 10 pairs: sw-02 (-0.6364), sw-04 (-0.0909).

Drew on 5 of 10 pairs: sw-01, sw-05, sw-06, sw-08, sw-10.

GPT-5 mini, Blind comparison of the same task, run twice, with ties counting half

Lost on 3 of 10 pairs: sw-01 (-1.0000), sw-05 (-1.0000), sw-06 (-1.0000).

Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key

Lost on 1 of 10 pairs: sw-03 (-0.0909).

Drew on 9 of 10 pairs: sw-01, sw-02, sw-04, sw-05, sw-06, sw-07, sw-08, sw-09, sw-10.

GPT-5 mini, Deterministic, scored against a committed answer key

Drew on 7 of 10 pairs: sw-01, sw-02, sw-03, sw-04, sw-06, sw-07, sw-09.

In Claude Code

These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument: no figure here is averaged with one tested via API above. How the two were compared.

  • Claude Fable 5 · unclear, and both runs agree
  • Claude Opus 5 · unclear, and both runs agree
  • Claude Haiku 4.5 · unclear under repeat sampling
  • Claude Sonnet 5 · unclear under repeat sampling
  • Claude Fable 5.1 · unclear in the version re-test

Skill cohort (one pass)

one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 50.92730.9364+0.009-0.023 to 0.04110In free drift1,533
Claude Opus 50.91820.9182+0.000-0.027 to 0.02710In free drift1,533
Thinking tokens, recorded per arm. Reported, never scored: no interval or Orbit on this page consults them.
ModelArmRecords that reasonedMean thinking tokensMost on one record
Claude Fable 5with skill20 of 2043106
Claude Fable 5no skill20 of 202037

Panel v2 (repeat sampling)

two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 50.96820.9273-0.041-0.066 to -0.01610In free drift46,092
Claude Haiku 4.50.97730.9682-0.009-0.035 to 0.01710In free drift32,024
Claude Opus 50.93180.9227-0.009-0.035 to 0.01710In free drift46,092
Claude Sonnet 50.94550.9409-0.004-0.029 to 0.02010In free drift59,132
Thinking tokens, recorded per arm. Reported, never scored: no interval or Orbit on this page consults them.
ModelArmRecords that reasonedMean thinking tokensMost on one record
Claude Fable 5with skill20 of 2043106
Claude Fable 5no skill20 of 202037

Fable 5.1, one pass

one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 5.10.97270.9545-0.018-0.054 to 0.01720In free driftnot measured

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 5.10.96360.9727+0.009-0.023 to 0.04110In free driftnot measured: this run records the dose per arm, in the doses list
Claude Haiku 4.50.97270.9636-0.009-0.041 to 0.02310In free driftnot measured: this run records the dose per arm, in the doses list
Claude Opus 50.91820.9364+0.018-0.017 to 0.05410In free driftnot measured: this run records the dose per arm, in the doses list
Claude Sonnet 50.94550.9455+0.000-0.027 to 0.02710In free driftnot measured: this run records the dose per arm, in the doses list

Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.

720 run records behind this page. Every verdict is computed at build time by the same frozen status_v1 rule that decides every other verdict on this site, and nothing here is written by hand.