Claude skills for SEO, tested
Skills for SEO work, measured on on-page audit tasks. Each was the same task, run twice, once with the skill loaded and once without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.
Measured on 3 models (Claude, GPT & Gemini): Claude Haiku 4.5, GPT-5 mini, Gemini 3.1 Flash Lite.
10 fixture pages seeded with 43 known on-page SEO issues from a closed list of 19 issue types. The model is asked to list the issues; the score is the share of seeded issues it names. 39 verified mechanically, 4 judgement items listed as such.
What helped, and on how many models
First
skills/seo-onpage rampstackco/claude-skills*
On-page audit
Helped on Gemini 3.1 Flash Lite; no reliable difference on Claude Haiku 4.5 and GPT-5 mini. Measured against no skill loaded.
Found 56% of the seeded on-page issues, against 39% with no skill loaded. Instruction arm: not measured.
Via API
- Claude Haiku 4.5 · unclear
- Gemini 3.1 Flash Lite · helped
- GPT-5 mini · unclear
In Claude Code
- Claude Fable 5, First run, single attempt · unclear
- Claude Opus 5, First run, single attempt · unclear
- Claude Fable 5, Second run, every task twice · unclear
- Claude Haiku 4.5, Second run, every task twice · unclear
- Claude Opus 5, Second run, every task twice · unclear
- Claude Sonnet 5, Second run, every task twice · unclear
- Claude Fable 5.1, Fable 5.1, single attempt · unclear
- Claude Fable 5.1, Wider run, three ways · unclear
- Claude Haiku 4.5, Wider run, three ways · unclear
- Claude Opus 5, Wider run, three ways · unclear
- Claude Sonnet 5, Wider run, three ways · unclear
Cost per task, with the skill loaded: about $0.00337. Adds about 6,200 tokens of context to every request.
Tied for second · shared with 3 others, and the tie is broken by nothing
skills/seo-onpage rampstackco/claude-skills*
On-page audit
No reliable difference on any model tested. Measured against the same task with a one-line instruction.
Found 63% of the seeded on-page issues, against 53% with no skill loaded. Against 53% with a one-line instruction instead.
Via API
- Gemini 3.1 Flash Lite · unclear
- GPT-5 mini · unclear
In Claude Code
- Claude Fable 5, First run, single attempt · unclear
- Claude Opus 5, First run, single attempt · unclear
- Claude Fable 5, Second run, every task twice · unclear
- Claude Haiku 4.5, Second run, every task twice · unclear
- Claude Opus 5, Second run, every task twice · unclear
- Claude Sonnet 5, Second run, every task twice · unclear
- Claude Fable 5.1, Fable 5.1, single attempt · unclear
- Claude Fable 5.1, Wider run, three ways · unclear
- Claude Haiku 4.5, Wider run, three ways · unclear
- Claude Opus 5, Wider run, three ways · unclear
- Claude Sonnet 5, Wider run, three ways · unclear
Cost per task, with the skill loaded: about $0.00161. Adds about 6,000 tokens of context to every request.
Tied for second · shared with 3 others, and the tie is broken by nothing
skills/seo affaan-m/ECC
On-page audit
No reliable difference on any model tested. Measured against no skill loaded.
Found 44% of the seeded on-page issues, against 39% with no skill loaded. Instruction arm: not measured.
Via API
- Claude Haiku 4.5 · unclear
- Gemini 3.1 Flash Lite · unclear
- GPT-5 mini · unclear
In Claude Code
- Claude Fable 5, First run, single attempt · unclear
- Claude Opus 5, First run, single attempt · unclear
- Claude Fable 5, Second run, every task twice · unclear
- Claude Haiku 4.5, Second run, every task twice · unclear
- Claude Opus 5, Second run, every task twice · unclear
- Claude Sonnet 5, Second run, every task twice · unclear
- Claude Fable 5.1, Fable 5.1, single attempt · unclear
- Claude Fable 5.1, Wider run, three ways · unclear
- Claude Haiku 4.5, Wider run, three ways · unclear
- Claude Opus 5, Wider run, three ways · unclear
- Claude Sonnet 5, Wider run, three ways · unclear
Cost per task, with the skill loaded: about $0.00102. Adds about 1,600 tokens of context to every request.
Tied for second · shared with 3 others, and the tie is broken by nothing
marketing-skill/skills/seo-audit alirezarezvani/claude-skills
On-page audit
No reliable difference on any model tested. Measured against the same task with a one-line instruction.
Found 58% of the seeded on-page issues, against 47% with no skill loaded. Against 56% with a one-line instruction instead.
Via API
- Gemini 3.1 Flash Lite · unclear
- GPT-5 mini · unclear
In Claude Code
- Claude Fable 5.1, Wider run, three ways · unclear
- Claude Haiku 4.5, Wider run, three ways · unclear
- Claude Opus 5, Wider run, three ways · unclear
- Claude Sonnet 5, Wider run, three ways · unclear
Cost per task, with the skill loaded: about $0.00153. Adds about 5,600 tokens of context to every request.
Tied for second · shared with 3 others, and the tie is broken by nothing
skills/seo affaan-m/ECC
On-page audit
No reliable difference on any model tested. Measured against the same task with a one-line instruction.
Found 56% of the seeded on-page issues, against 51% with no skill loaded. Against 56% with a one-line instruction instead.
Via API
- Gemini 3.1 Flash Lite · unclear
- GPT-5 mini · unclear
In Claude Code
- Claude Fable 5, First run, single attempt · unclear
- Claude Opus 5, First run, single attempt · unclear
- Claude Fable 5, Second run, every task twice · unclear
- Claude Haiku 4.5, Second run, every task twice · unclear
- Claude Opus 5, Second run, every task twice · unclear
- Claude Sonnet 5, Second run, every task twice · unclear
- Claude Fable 5.1, Fable 5.1, single attempt · unclear
- Claude Fable 5.1, Wider run, three ways · unclear
- Claude Haiku 4.5, Wider run, three ways · unclear
- Claude Opus 5, Wider run, three ways · unclear
- Claude Sonnet 5, Wider run, three ways · unclear
Cost per task, with the skill loaded: about $0.000505. Adds about 1,600 tokens of context to every request.
The numbers
How to read these tables
How to read these tables
- Without -> with
- The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
- Effect (range)
- With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
- Orbit
- Stable: delta at or above the pass threshold, and the interval excludes zero
- Past the horizon: delta at or below the failure floor, and the interval excludes zero
- In free drift: the interval spans zero, or the delta sits between the floor and the threshold
- Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound
A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.
- Cohort
- The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
- Commit
- The exact version tested. Later commits are not covered by this verdict.
- Injected context tokens
- How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
- Tested
- The date. Verdicts age as models change.
On-page audit
| Skill | Cohort | Without -> with | Effect (range) | Orbit | Injected context tokens | Tested |
|---|---|---|---|---|---|---|
| No skill loaded First run, with and without the skill | baseline | 39% | 0 by definition | 0 | 2026-08-30 | |
| One-line instruction instead First run, with and without the skill | baseline | instruction arm: not measured | 0 by definition | 0 | 2026-08-30 | |
skills/seo-onpage from rampstackco/claude-skills 047924252254 First run, with and without the skill Best result in this test | rampstackco/claude-skills* | 39% -> 56% | +0.17 [+0.04, +0.23] | In free drift on 2 of 3 | 6,184 | 2026-08-30 |
skills/seo from affaan-m/ECC d8409a4b0813 First run, with and without the skill | affaan-m/ECC | 39% -> 44% | +0.05 [+0.00, +0.12] | In free drift | 1,637 | 2026-08-30 |
| No skill loaded Wider run, three ways | baseline | 51% | 0 by definition | 0 | 2026-09-05 | |
| One-line instruction instead Wider run, three ways | baseline | 55% | 0 by definition | 0 | 2026-09-05 | |
skills/seo-onpage from rampstackco/claude-skills 047924252254 Wider run, three ways Best result in this test | rampstackco/claude-skills* | 53% -> 63% | +0.10 [+0.08, +0.13] | In free drift | 5,975 | 2026-09-05 |
marketing-skill/skills/seo-audit from alirezarezvani/claude-skills 19392f7a0826 Wider run, three ways | alirezarezvani/claude-skills | 47% -> 58% | +0.02 [-0.04, +0.07] | In free drift | 5,625 | 2026-09-05 |
skills/seo from affaan-m/ECC d8409a4b0813 Wider run, three ways | affaan-m/ECC | 51% -> 56% | -0.00 [-0.06, +0.05] | In free drift | 1,578 | 2026-09-05 |
| Averaged over three models; per-model results on each skill's page. Ranked within this class only. | ||||||
Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.