skills/skill-creator from anthropics/skills
Skill authoring tasks, run with the skill loaded and without it, on the same task pairs.
Identity pin
- Repository
- anthropics/skills
- Path
- skills/skill-creator
- Commit
3b3fad96af16a10759d930941b4520ba0c40edae- Content hash
e51ce07ea7fdd8aa453aa9c03e123faa44c6c810f7e3935097124674be29b734- Date tested
- not yet tested
The pin is the whole of this skill's identity here. It resolves at https://github.com/anthropics/skills/tree/3b3fad96af16a10759d930941b4520ba0c40edae/skills/skill-creator, and the content hash is a sha256 over exactly the text the with arm was given. Nothing else about the skill appears on this site.
What was claimed
- Loading skills/skill-creator from anthropics/skills improves conformance to the open Agent Skills specification on skill authoring tasks.
Pass criterion: With-arm mean conformance to the open Agent Skills specification exceeds the without-arm by at least 0.10 on the same 10 briefs at the same model and settings. Key B, the house standard, is reported beside this figure and takes no part in it. Scale: unit. Fraction of the applicable Agent Skills specification predicates passed, so a [0,1] fraction, which is the unit scale. The 0.2 unit threshold in status_v1 applies unchanged. A predicate for an optional field the document does not offer is not applicable and leaves the denominator, so a document is never credited for a rule it did not engage.
TWO KEYS, ONE RANKED. Key A is the open Agent Skills specification, pinned at harness/artifacts/keys/agentskills-specification. It is published by neither cohort, it is the format both skills in this class claim to teach, and it is the only key that becomes a verdict. Key B is SKILL_AUTHORING.md, the operator's own house standard: it is still computed and shown, under a disclosure naming whose standard it is, and it is excluded from the cohort ranking and from every cross-cohort statement. The earlier version of this class ranked on Key B alone, which made a lift here partly a measure of agreement with the operator's conventions. Neither key is given to either arm, and the task prompt restates no part of either.
Injected context tokens
| Arm | Context characters | Injected context tokens |
|---|---|---|
| without | 0 | 0 |
| with | 45,101 | 10,587 |
The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.
Not yet measured
No run records exist for this skill. The instrument is built and committed, the task sets are frozen, and the projected cost of the run is published, but no model has been called. There is therefore no verdict, no effect, no interval and no cost per task on this page, and none is shown.
- S02-skill-creator: 10 task pairs per model. Deterministic, scored against a committed answer key.
Key A: the open Agent Skills specification
This class is ranked on conformance to the open specification and on nothing else. The specification is published by neither cohort, it is given to neither arm, and it is pinned like any other artifact: agentskills/agentskills docs at 69ef37e9424c0a7ea9dd2293b559e43ec8176379, content hash b9079c0c10b7930e8c6a20ff2bc10cda2a3343c55185120e3f1116a1a529b220.
| Predicate | What it requires, in our words | Where the specification says it | Scored when |
|---|---|---|---|
frontmatter-present | The document opens with a YAML frontmatter block, and Markdown follows it. | SKILL.md format, opening sentence | always |
frontmatter-name-present | A name field is present. | SKILL.md format, frontmatter table, name row | always |
frontmatter-description-present | A description field is present. | SKILL.md format, frontmatter table, description row | always |
name-within-length-limit | The name is between one and sixty four characters. | name field, first bullet | always |
name-characters-legal | The name uses only lowercase letters, digits and hyphens. | name field, second bullet | always |
name-no-edge-hyphen | The name neither opens nor closes on a hyphen. | name field, third bullet | always |
name-no-consecutive-hyphens | The name carries no run of two hyphens. | name field, fourth bullet | always |
description-within-length-limit | The description is between one and one thousand and twenty four characters. | description field, first bullet | always |
body-present | Instructions follow the frontmatter rather than an empty document. | Body content, opening sentence | always |
within-line-limit | The document stays under the stated line ceiling for a main SKILL.md. | Progressive disclosure, closing line | always |
file-references-one-level-deep | Every relative path the document points at sits at most one directory down. | File references, closing line | always |
compatibility-within-length-limit | If compatibility is given, it stays inside its stated ceiling. | compatibility field, first bullet | the field is offered |
metadata-is-string-map | If metadata is given, it maps string keys to string values. | metadata field, first bullet | the field is offered |
allowed-tools-is-one-string | If allowed-tools is given, it is a single space separated string. | allowed-tools field, first bullet | the field is offered |
A predicate for an optional field a document does not offer is not applicable and leaves the denominator, so a document is never credited for a rule it did not engage.
Key B: the operator's house standard, not ranked
No records exist, so Key B has no figures either. It is described here rather than left out, because a reader deciding what the ranked number means needs to know what the unranked one is.
No interval, no orbit and no verdict appears above, and none can: these are two arm means and the frozen rule is never called on them.
The committed instrument is at harness/claims/claims-skills.json, harness/claims/tasksets-skills.json and harness/claims/skill-cohort.json. The projection for the run that has not happened is at harness/results/dry-run-projection-skills.md. How a claim gets tested.