Does Skill creation walkthrough help? Tested on writing skills tasks

Skill creation walkthrough, from rampstackco/claude-skills. Its skill authoring tasks were run with the skill loaded and without it. Each time it is the same task, run twice.

What we tested

Whether loading this skill helps a model write a skill document.

What counts as helping

The score with the skill has to beat the score without it by at least 10 percentage points, on the same 10 briefs.

How we scored it

We checked each document against the open Agent Skills specification and counted the rules it passed. The score is the share it passed, from 0 to 1.

Without to with, per model

Via API tested via API

Claude Haiku 4.5 First run, with and without the skill

0.618
1.000

Holds

Gemini 3.1 Flash Lite First run, with and without the skill

0.873
1.000

No measured effect

Gemini 3.1 Flash Lite Wider run, three ways

0.745
1.000

Holds

GPT-5 mini First run, with and without the skill

0.718
1.000

Holds

GPT-5 mini Wider run, three ways

0.618
1.000

Holds

deterministic pass rate, 0 to 1

In Claude Code tested in Claude Code

Claude Fable 5.1 Fable 5.1, single attempt

1.000 without and with

Could not measure

Claude Fable 5 Second run, every task twice

1.000
0.927

No measured effect

Claude Fable 5 First run, single attempt

1.000
0.927

No measured effect

Claude Fable 5.1 Wider run, three ways

1.000
0.855

No measured effect

Claude Haiku 4.5 Second run, every task twice

0.691
1.000

Holds

Claude Haiku 4.5 Wider run, three ways

0.545
1.000

Holds

Claude Opus 5 First run, single attempt

1.000 without and with

Could not measure

Claude Opus 5 Second run, every task twice

1.000 without and with

Could not measure

Claude Opus 5 Wider run, three ways

1.000 without and with

Could not measure

Claude Sonnet 5 Second run, every task twice

1.000
0.891

No measured effect

Claude Sonnet 5 Wider run, three ways

1.000
0.855

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Grey is the score without. Colour is the score with. The bracket shows how much the difference could move if we ran it again. Rows are ordered by the score with. A model measured at two versions keeps its versions next to each other. The two test methods are reported separately and never averaged. Where the two scores are the same, the number is printed once at the end of the pair. A bracket wider than the axis is drawn to the edge with its cap omitted; the table gives its bounds.
Show per-task detail

Every item, without and with, per model

Via API tested via API

  • Claude Haiku 4.5
  • GPT-5 mini
  • GPT-5 mini
  • Gemini 3.1 Flash Lite
  • Gemini 3.1 Flash Lite

deterministic pass rate, 0 to 1

Hover or focus an item to read its task and every model’s two values.

Task-paired: these item means are what this cell's interval is built from.

The item means behind this plot
Per-item arm means. Without is the unaided arm, with is the treated arm.
ItemClaude Haiku 4.5 GPT-5 mini GPT-5 mini Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite
Task without with without with without with without with without with
sa-05 0.545 1.000 0.273 1.000 0.273 1.000 0.545 1.000 0.909 1.000
sa-10 0.545 1.000 0.727 1.000 0.727 1.000 0.545 1.000 0.545 1.000
sa-02 0.545 1.000 0.273 1.000 0.818 1.000 0.909 1.000 0.909 1.000
sa-08 0.545 1.000 0.909 1.000 0.727 1.000 0.364 1.000 0.909 1.000
sa-07 0.545 1.000 0.273 1.000 0.909 1.000 0.909 1.000 0.909 1.000
sa-09 0.545 1.000 0.909 1.000 0.273 1.000 0.909 1.000 0.909 1.000
sa-06 0.545 1.000 0.727 1.000 0.909 1.000 0.545 1.000 0.909 1.000
sa-04 0.909 1.000 0.273 1.000 0.909 1.000 0.909 1.000 0.909 1.000
sa-01 0.545 1.000 0.909 1.000 0.727 1.000 0.909 1.000 0.909 1.000
sa-03 0.909 1.000 0.909 1.000 0.909 1.000 0.909 1.000 0.909 1.000
A picture of the per-item means behind the per-model numbers. The tables are the source. Each model draws two lines over the same task set: a dashed line through its unaided scores and a solid line through its treated ones. Items are ordered by the mean unaided score across the models that measured them, lowest first. The two instruments are reported separately and never averaged.

Without to with, per model

In Claude Code tested in Claude Code

Claude Haiku 4.5 Expansion cohort (three arms)

Holds

Claude Haiku 4.5 Panel v2 (repeat sampling)

Holds

Claude Fable 5 Panel v2 (repeat sampling)

No measured effect

Claude Fable 5 Skill cohort (one pass)

No measured effect

Claude Fable 5.1 Expansion cohort (three arms)

No measured effect

Claude Fable 5.1 Fable 5.1, one pass

Could not measure

Claude Opus 5 Expansion cohort (three arms)

Could not measure

Claude Opus 5 Panel v2 (repeat sampling)

Could not measure

Claude Opus 5 Skill cohort (one pass)

Could not measure

Claude Sonnet 5 Expansion cohort (three arms)

No measured effect

Claude Sonnet 5 Panel v2 (repeat sampling)

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Each row runs from the score without to the score with. The thin bar beneath it is the bracket, and it shows how much the difference could move if we ran it again. The two test methods are reported separately and never averaged. A bracket wider than the axis is drawn to the edge with its end cap omitted; the table gives its bounds.

The Claude Code runs committed their cells as arm means and intervals rather than as per-item values, so this panel is drawn without to with rather than item by item.

What exactly was tested, and how it was scored
Repository
rampstackco/claude-skills
Path
skills/skill-creation-walkthrough
Commit
0479242522549dfdb389bb9b7807ad4d6016ffb7
Content hash
8d539bb7a97786875b4489a9005434c2f7c772c8f009994a5ae1ef2b66e8da80
Date tested
2026-08-30

The pin is the whole of this skill's identity here. It resolves at https://github.com/rampstackco/claude-skills/tree/0479242522549dfdb389bb9b7807ad4d6016ffb7/skills/skill-creation-walkthrough, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.

S01-skill-creation-walkthrough

Loading skills/skill-creation-walkthrough from rampstackco/claude-skills improves conformance to the open Agent Skills specification on skill authoring tasks.

Pass criterion
With-arm mean conformance to the open Agent Skills specification exceeds the without-arm by at least 0.10 on the same 10 briefs at the same model and settings. Key B, the house standard, is reported beside this figure and takes no part in it.
Scale
unit. Fraction of the applicable Agent Skills specification predicates passed, so a [0,1] fraction, which is the unit scale. The 0.2 unit threshold in status_v1 applies unchanged. A predicate for an optional field the document does not offer is not applicable and leaves the denominator, so a document is never credited for a rule it did not engage.
Task pairs planned per model
10
Notes
TWO KEYS, ONE RANKED. Key A is the open Agent Skills specification, pinned at harness/artifacts/keys/agentskills-specification. It is published by neither cohort, it is the format both skills in this class claim to teach, and it is the only key that becomes a verdict. Key B is SKILL_AUTHORING.md, the operator's own house standard: it is still computed and shown, under a disclosure naming whose standard it is, and it is excluded from the cohort ranking and from every cross-cohort statement. The earlier version of this class ranked on Key B alone, which made a lift here partly a measure of agreement with the operator's conventions. Neither key is given to either arm, and the task prompt restates no part of either.

Injected context tokens

Injected context tokens, per arm
ArmContext charactersInjected context tokens
without00
with37,0088,687

The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.

Verdict per model

Effect per model
ModelReadingOrbitWithout -> withEffect95 percent intervalTask pairsModel version returned
Claude Haiku 4.5Deterministic, scored against a committed answer keyStable0.6182 -> 1.0000+0.3818[0.2868, 0.4768]10 pairs, 20 of 20 recordsclaude-haiku-4-5-20251001
Gemini 3.1 Flash LiteDeterministic, scored against a committed answer keyIn free drift unclear. The model scored 87% without it.0.8727 -> 1.0000+0.1273[0.0560, 0.1985]10 pairs, 20 of 20 recordsgemini-3.1-flash-lite
GPT-5 miniDeterministic, scored against a committed answer keyStable0.7182 -> 1.0000+0.2818[0.1282, 0.4354]10 pairs, 20 of 20 recordsgpt-5-mini-2025-08-07
Gemini 3.1 Flash LiteDeterministic, scored against a committed answer keyStable0.7455 -> 1.0000+0.2545[0.1196, 0.3895]10 pairs, 30 of 30 recordsgemini-3.1-flash-lite
GPT-5 miniDeterministic, scored against a committed answer keyStable (10 of 10 pairs)0.6182 -> 1.0000+0.3818[0.1925, 0.5711]10 pairs, 30 of 33 recordsgpt-5-mini-2025-08-07
  • Stable: the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero.
  • In free drift: no separation the design can resolve. Not evidence of no effect.

Cost

Cost per task
ModelArmTasks attemptedMean input tokensMean output tokensCost per taskCost ratioGrading cost per task pair
Claude Haiku 4.5 served claude-haiku-4-5-20251001without10681,210$0.00612
Claude Haiku 4.5 served claude-haiku-4-5-20251001with108,7911,533$0.016452.69x
GPT-5 mini served gpt-5-mini-2025-08-07without10611,459$0.00293
GPT-5 mini served gpt-5-mini-2025-08-07with108,0031,583$0.005171.76x
Gemini 3.1 Flash Litewithout1059580$0.00088
Gemini 3.1 Flash Litewith108,376713$0.003163.58x

Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.

In Claude Code

These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument: no figure here is averaged with one tested via API above. How the two were compared.

  • Claude Fable 5 · unclear, and both runs agree
  • Claude Opus 5 · could not be measured, and both runs agree
  • Claude Haiku 4.5 · helped under repeat sampling
  • Claude Sonnet 5 · unclear under repeat sampling
  • Claude Fable 5.1 · could not be measured in the version re-test

Skill cohort (one pass)

one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 51.00000.9273-0.073-0.215 to 0.07010In free drift12,321
Claude Opus 51.00001.0000+0.0000.000 to 0.00010Unobservable12,321
Thinking tokens, recorded per arm. Reported, never scored: no interval or Orbit on this page consults them.
ModelArmRecords that reasonedMean thinking tokensMost on one record
Claude Fable 5with skill20 of 2068231
Claude Fable 5no skill20 of 2064127

Panel v2 (repeat sampling)

two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 51.00000.9273-0.073-0.168 to 0.02210In free drift258,944
Claude Haiku 4.50.69091.0000+0.3090.220 to 0.39810Stable183,796
Claude Opus 51.00001.0000+0.0000.000 to 0.00010Unobservable258,944
Claude Sonnet 51.00000.8909-0.109-0.218 to -0.00010In free drift271,984
Thinking tokens, recorded per arm. Reported, never scored: no interval or Orbit on this page consults them.
ModelArmRecords that reasonedMean thinking tokensMost on one record
Claude Fable 5with skill20 of 2068231
Claude Fable 5no skill20 of 2064127

Fable 5.1, one pass

one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 5.11.00001.0000+0.0000.000 to 0.00020Unobservablenot measured

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokens
Claude Fable 5.11.00000.8545-0.145-0.336 to 0.04510In free driftnot measured: this run records the dose per arm, in the doses list
Claude Haiku 4.50.54551.0000+0.4550.455 to 0.45510Stablenot measured: this run records the dose per arm, in the doses list
Claude Opus 51.00001.0000+0.0000.000 to 0.00010Unobservablenot measured: this run records the dose per arm, in the doses list
Claude Sonnet 51.00000.8545-0.145-0.336 to 0.04510In free driftnot measured: this run records the dose per arm, in the doses list

Key A: the open Agent Skills specification

This class is ranked on conformance to the open specification and on nothing else. The specification is published by neither cohort, it is given to neither arm, and it is pinned like any other artifact: agentskills/agentskills docs at 69ef37e9424c0a7ea9dd2293b559e43ec8176379, content hash b9079c0c10b7930e8c6a20ff2bc10cda2a3343c55185120e3f1116a1a529b220.

Key A predicates, each with the passage it comes from
PredicateWhat it requires, in our wordsWhere the specification says itScored when
frontmatter-presentThe document opens with a YAML frontmatter block, and Markdown follows it.SKILL.md format, opening sentencealways
frontmatter-name-presentA name field is present.SKILL.md format, frontmatter table, name rowalways
frontmatter-description-presentA description field is present.SKILL.md format, frontmatter table, description rowalways
name-within-length-limitThe name is between one and sixty four characters.name field, first bulletalways
name-characters-legalThe name uses only lowercase letters, digits and hyphens.name field, second bulletalways
name-no-edge-hyphenThe name neither opens nor closes on a hyphen.name field, third bulletalways
name-no-consecutive-hyphensThe name carries no run of two hyphens.name field, fourth bulletalways
description-within-length-limitThe description is between one and one thousand and twenty four characters.description field, first bulletalways
body-presentInstructions follow the frontmatter rather than an empty document.Body content, opening sentencealways
within-line-limitThe document stays under the stated line ceiling for a main SKILL.md.Progressive disclosure, closing linealways
file-references-one-level-deepEvery relative path the document points at sits at most one directory down.File references, closing linealways
compatibility-within-length-limitIf compatibility is given, it stays inside its stated ceiling.compatibility field, first bulletthe field is offered
metadata-is-string-mapIf metadata is given, it maps string keys to string values.metadata field, first bulletthe field is offered
allowed-tools-is-one-stringIf allowed-tools is given, it is a single space separated string.allowed-tools field, first bulletthe field is offered

A predicate for an optional field a document does not offer is not applicable and leaves the denominator, so a document is never credited for a rule it did not engage.

Key B: the operator's house standard, not ranked

Key B: the operator's house standard, not ranked
ClaimModelWithoutWithTask pairs
S01-skill-creation-walkthroughClaude Haiku 4.50.25830.475010
S01-skill-creation-walkthroughGPT-5 mini0.20830.375010
S01-skill-creation-walkthroughGemini 3.1 Flash Lite0.25000.500010

No interval, no orbit and no verdict appears above, and none can: these are two arm means and the frozen rule is never called on them.

Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.

720 run records behind this page. Every verdict is computed at build time by the same frozen status_v1 rule that decides every other verdict on this site, and nothing here is written by hand.