Does Skill creator (anthropics) help? Tested on writing skills tasks
Skill creator (anthropics), from anthropics/skills. Its skill authoring tasks were run with the skill loaded and without it. Each time it is the same task, run twice.
What we tested
Whether loading this skill helps a model write a skill document.
What counts as helping
The score with the skill has to beat the score without it by at least 10 percentage points, on the same 10 briefs.
How we scored it
We checked each document against the open Agent Skills specification and counted the rules it passed. The score is the share it passed, from 0 to 1.
Without to with, per model
Via API tested via API
Claude Haiku 4.5 First run, with and without the skill
Holds
Gemini 3.1 Flash Lite First run, with and without the skill
No measured effect
Gemini 3.1 Flash Lite Wider run, three ways
Holds
GPT-5 mini First run, with and without the skill
Holds
GPT-5 mini Wider run, three ways
Holds
deterministic pass rate, 0 to 1
In Claude Code tested in Claude Code
Claude Fable 5 First run, single attempt
Could not measure
Claude Fable 5 Second run, every task twice
Could not measure
Claude Fable 5.1 Fable 5.1, single attempt
Scored worse
Claude Fable 5.1 Wider run, three ways
Scored worse
Claude Haiku 4.5 Second run, every task twice
Holds
Claude Haiku 4.5 Wider run, three ways
Holds
Claude Opus 5 First run, single attempt
Could not measure
Claude Opus 5 Wider run, three ways
Could not measure
Claude Opus 5 Second run, every task twice
No measured effect
Claude Sonnet 5 Wider run, three ways
Could not measure
Claude Sonnet 5 Second run, every task twice
No measured effect
deterministic pass rate, 0 to 1
Show per-task detail
Every item, without and with, per model
Via API tested via API
- Claude Haiku 4.5
- GPT-5 mini
- GPT-5 mini
- Gemini 3.1 Flash Lite
- Gemini 3.1 Flash Lite
deterministic pass rate, 0 to 1
Hover or focus an item to read its task and every model’s two values.
Task-paired: these item means are what this cell's interval is built from.
The item means behind this plot
| Item | Claude Haiku 4.5 | GPT-5 mini | GPT-5 mini | Gemini 3.1 Flash Lite | Gemini 3.1 Flash Lite | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | without | with | without | with | without | with | without | with | without | with |
| sa-08 | 0.545 | 1.000 | 0.545 | 1.000 | not measured | not measured | 0.364 | 1.000 | 0.909 | 1.000 |
| sa-10 | 0.545 | 1.000 | 0.727 | 1.000 | not measured | not measured | 0.545 | 1.000 | 0.545 | 1.000 |
| sa-05 | 0.545 | 1.000 | 0.818 | 1.000 | 0.273 | 1.000 | 0.545 | 1.000 | 0.909 | 1.000 |
| sa-09 | 0.545 | 1.000 | 0.273 | 1.000 | 0.727 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 |
| sa-06 | 0.545 | 1.000 | 0.727 | 1.000 | not measured | not measured | 0.545 | 1.000 | 0.909 | 1.000 |
| sa-02 | 0.545 | 1.000 | 0.273 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 |
| sa-01 | 0.545 | 1.000 | 0.909 | 1.000 | not measured | not measured | 0.909 | 1.000 | 0.909 | 1.000 |
| sa-07 | 0.545 | 1.000 | 0.909 | 1.000 | 0.818 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 |
| sa-03 | 0.909 | 1.000 | 0.727 | 1.000 | 0.727 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 |
| sa-04 | 0.909 | 1.000 | 0.727 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 | 0.909 | 1.000 |
Without to with, per model
In Claude Code tested in Claude Code
Claude Haiku 4.5 Expansion cohort (three arms)
Holds
Claude Haiku 4.5 Panel v2 (repeat sampling)
Holds
Claude Fable 5 Panel v2 (repeat sampling)
Could not measure
Claude Fable 5 Skill cohort (one pass)
Could not measure
Claude Fable 5.1 Expansion cohort (three arms)
Scored worse
Claude Fable 5.1 Fable 5.1, one pass
Scored worse
Claude Opus 5 Expansion cohort (three arms)
Could not measure
Claude Opus 5 Panel v2 (repeat sampling)
No measured effect
Claude Opus 5 Skill cohort (one pass)
Could not measure
Claude Sonnet 5 Expansion cohort (three arms)
Could not measure
Claude Sonnet 5 Panel v2 (repeat sampling)
No measured effect
deterministic pass rate, 0 to 1
The Claude Code runs committed their cells as arm means and intervals rather than as per-item values, so this panel is drawn without to with rather than item by item.
What exactly was tested, and how it was scored
- Repository
- anthropics/skills
- Path
- skills/skill-creator
- Commit
3b3fad96af16a10759d930941b4520ba0c40edae- Content hash
e51ce07ea7fdd8aa453aa9c03e123faa44c6c810f7e3935097124674be29b734- Date tested
- 2026-08-30
The pin is the whole of this skill's identity here. It resolves at https://github.com/anthropics/skills/tree/3b3fad96af16a10759d930941b4520ba0c40edae/skills/skill-creator, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.
S02-skill-creator
Loading skills/skill-creator from anthropics/skills improves conformance to the open Agent Skills specification on skill authoring tasks.
- Pass criterion
- With-arm mean conformance to the open Agent Skills specification exceeds the without-arm by at least 0.10 on the same 10 briefs at the same model and settings. Key B, the house standard, is reported beside this figure and takes no part in it.
- Scale
- unit. Fraction of the applicable Agent Skills specification predicates passed, so a [0,1] fraction, which is the unit scale. The 0.2 unit threshold in status_v1 applies unchanged. A predicate for an optional field the document does not offer is not applicable and leaves the denominator, so a document is never credited for a rule it did not engage.
- Task pairs planned per model
- 10
- Notes
- TWO KEYS, ONE RANKED. Key A is the open Agent Skills specification, pinned at harness/artifacts/keys/agentskills-specification. It is published by neither cohort, it is the format both skills in this class claim to teach, and it is the only key that becomes a verdict. Key B is SKILL_AUTHORING.md, the operator's own house standard: it is still computed and shown, under a disclosure naming whose standard it is, and it is excluded from the cohort ranking and from every cross-cohort statement. The earlier version of this class ranked on Key B alone, which made a lift here partly a measure of agreement with the operator's conventions. Neither key is given to either arm, and the task prompt restates no part of either.
Injected context tokens
| Arm | Context characters | Injected context tokens |
|---|---|---|
| without | 0 | 0 |
| with | 45,101 | 10,587 |
The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.
Verdict per model
| Model | Reading | Orbit | Without -> with | Effect | 95 percent interval | Task pairs | Model version returned |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | Deterministic, scored against a committed answer key | Stable | 0.6182 -> 1.0000 | +0.3818 | [0.2868, 0.4768] | 10 pairs, 20 of 20 records | claude-haiku-4-5-20251001 |
| Gemini 3.1 Flash Lite | Deterministic, scored against a committed answer key | In free drift unclear. The model scored 87% without it. | 0.8727 -> 1.0000 | +0.1273 | [0.0560, 0.1985] | 10 pairs, 20 of 20 records | gemini-3.1-flash-lite |
| GPT-5 mini | Deterministic, scored against a committed answer key | Stable (6 of 10 pairs, all 4 lost on the with arm: sa-01, sa-06, sa-08, sa-10) | 0.7273 -> 1.0000 | +0.2727 | [0.0830, 0.4624] | 6 pairs, 16 of 20 records | gpt-5-mini-2025-08-07 |
| Gemini 3.1 Flash Lite | Deterministic, scored against a committed answer key | Stable | 0.7455 -> 1.0000 | +0.2545 | [0.1196, 0.3895] | 10 pairs, 30 of 30 records | gemini-3.1-flash-lite |
| GPT-5 mini | Deterministic, scored against a committed answer key | Stable (10 of 10 pairs) | 0.6636 -> 1.0000 | +0.3364 | [0.1932, 0.4795] | 10 pairs, 30 of 82 records | gpt-5-mini-2025-08-07, gpt-5-mini |
- Stable: the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero.
- In free drift: no separation the design can resolve. Not evidence of no effect.
Cost
| Model | Arm | Tasks attempted | Mean input tokens | Mean output tokens | Cost per task | Cost ratio | Grading cost per task pair |
|---|---|---|---|---|---|---|---|
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | without | 10 | 68 | 1,233 | $0.00623 | ||
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | with | 10 | 12,020 | 1,836 | $0.02120 | 3.40x | |
GPT-5 mini served gpt-5-mini-2025-08-07 | without | 10 | 61 | 1,497 | $0.00301 | ||
GPT-5 mini served gpt-5-mini-2025-08-07 | with | 10 | 16,040 | 2,947 | $0.00990 | 3.29x | |
| Gemini 3.1 Flash Lite | without | 10 | 59 | 580 | $0.00088 | ||
| Gemini 3.1 Flash Lite | with | 10 | 11,827 | 506 | $0.00372 | 4.20x |
Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.
In Claude Code
These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument: no figure here is averaged with one tested via API above. How the two were compared.
- Claude Fable 5 · could not be measured, and both runs agree
- Claude Opus 5 · unclear under repeat sampling; could not be measured in the single-pass run
- Claude Haiku 4.5 · helped under repeat sampling
- Claude Sonnet 5 · unclear under repeat sampling
- Claude Fable 5.1 · made results worse in the version re-test
Skill cohort (one pass)
one pass over each committed item. Matrix hash 2707a06a. No pre-registration: this run predates the practice on this arm. Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 1.0000 | 1.0000 | +0.000 | 0.000 to 0.000 | 10 | Unobservable | 15,961 |
| Claude Opus 5 | 1.0000 | 1.0000 | +0.000 | 0.000 to 0.000 | 10 | Unobservable | 15,961 |
| Model | Arm | Records that reasoned | Mean thinking tokens | Most on one record |
|---|---|---|---|---|
| Claude Fable 5 | with skill | 20 of 20 | 68 | 231 |
| Claude Fable 5 | no skill | 20 of 20 | 64 | 127 |
Panel v2 (repeat sampling)
two passes over each committed item, under repeat sampling. Matrix hash 94bab960. Pre-registration · Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 1.0000 | 1.0000 | +0.000 | 0.000 to 0.000 | 1 | Unobservable | 16,933 |
| Claude Haiku 4.5 | 0.7091 | 1.0000 | +0.291 | 0.227 to 0.355 | 10 | Stable | 248,376 |
| Claude Opus 5 | 1.0000 | 0.9091 | -0.091 | -0.269 to 0.087 | 4 | In free drift | 133,338 |
| Claude Sonnet 5 | 1.0000 | 0.9273 | -0.073 | -0.168 to 0.022 | 10 | In free drift | 344,784 |
| Model | Arm | Records that reasoned | Mean thinking tokens | Most on one record |
|---|---|---|---|---|
| Claude Fable 5 | with skill | 20 of 20 | 68 | 231 |
| Claude Fable 5 | no skill | 20 of 20 | 64 | 127 |
Fable 5.1, one pass
one pass over each committed item. Matrix hash db373661. Pre-registration · Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 1.0000 | 0.7818 | -0.218 | -0.436 to -0.000 | 20 | Past the horizon | not measured |
Expansion cohort (three arms)
one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.
| Model | No skill | With skill | Delta | Interval | Pairs | Orbit | Injected tokens |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 1.0000 | 0.5636 | -0.436 | -0.669 to -0.204 | 10 | Past the horizon | not measured: this run records the dose per arm, in the doses list |
| Claude Haiku 4.5 | 0.5636 | 1.0000 | +0.436 | 0.353 to 0.520 | 10 | Stable | not measured: this run records the dose per arm, in the doses list |
| Claude Opus 5 | 1.0000 | 1.0000 | +0.000 | 0.000 to 0.000 | 10 | Unobservable | not measured: this run records the dose per arm, in the doses list |
| Claude Sonnet 5 | 1.0000 | 1.0000 | +0.000 | 0.000 to 0.000 | 10 | Unobservable | not measured: this run records the dose per arm, in the doses list |
Key A: the open Agent Skills specification
This class is ranked on conformance to the open specification and on nothing else. The specification is published by neither cohort, it is given to neither arm, and it is pinned like any other artifact: agentskills/agentskills docs at 69ef37e9424c0a7ea9dd2293b559e43ec8176379, content hash b9079c0c10b7930e8c6a20ff2bc10cda2a3343c55185120e3f1116a1a529b220.
| Predicate | What it requires, in our words | Where the specification says it | Scored when |
|---|---|---|---|
frontmatter-present | The document opens with a YAML frontmatter block, and Markdown follows it. | SKILL.md format, opening sentence | always |
frontmatter-name-present | A name field is present. | SKILL.md format, frontmatter table, name row | always |
frontmatter-description-present | A description field is present. | SKILL.md format, frontmatter table, description row | always |
name-within-length-limit | The name is between one and sixty four characters. | name field, first bullet | always |
name-characters-legal | The name uses only lowercase letters, digits and hyphens. | name field, second bullet | always |
name-no-edge-hyphen | The name neither opens nor closes on a hyphen. | name field, third bullet | always |
name-no-consecutive-hyphens | The name carries no run of two hyphens. | name field, fourth bullet | always |
description-within-length-limit | The description is between one and one thousand and twenty four characters. | description field, first bullet | always |
body-present | Instructions follow the frontmatter rather than an empty document. | Body content, opening sentence | always |
within-line-limit | The document stays under the stated line ceiling for a main SKILL.md. | Progressive disclosure, closing line | always |
file-references-one-level-deep | Every relative path the document points at sits at most one directory down. | File references, closing line | always |
compatibility-within-length-limit | If compatibility is given, it stays inside its stated ceiling. | compatibility field, first bullet | the field is offered |
metadata-is-string-map | If metadata is given, it maps string keys to string values. | metadata field, first bullet | the field is offered |
allowed-tools-is-one-string | If allowed-tools is given, it is a single space separated string. | allowed-tools field, first bullet | the field is offered |
A predicate for an optional field a document does not offer is not applicable and leaves the denominator, so a document is never credited for a rule it did not engage.
Key B: the operator's house standard, not ranked
| Claim | Model | Without | With | Task pairs |
|---|---|---|---|---|
| S02-skill-creator | Claude Haiku 4.5 | 0.2583 | 0.1833 | 10 |
| S02-skill-creator | GPT-5 mini | 0.1944 | 0.1944 | 6 |
| S02-skill-creator | Gemini 3.1 Flash Lite | 0.2500 | 0.2500 | 10 |
No interval, no orbit and no verdict appears above, and none can: these are two arm means and the frozen rule is never called on them.
Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.