Claude skills for Brand Voice, tested

Skills for brand voice work, measured on voice tasks. Each was the same task, run twice, once with the skill loaded and once without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.

Measured on 3 models (Claude, GPT & Gemini): Claude Haiku 4.5, GPT-5 mini, Gemini 3.1 Flash Lite.

The model is given a brand brief and asked for a short piece of copy. There is no answer key; a grader sees the with and without outputs, blind and in random order, and picks one. The score is how often the skill's output was preferred.

What helped, and on how many models

  • First

    skills/brand-voice affaan-m/ECC

    Voice

    Helped clearly on Claude Haiku 4.5, Gemini 3.1 Flash Lite and GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 24 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · helped
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · helped

    Cost per task, with the skill loaded: about $0.00101. Adds about 1,200 tokens of context to every request.

  • Second

    skills/brand-voice rampstackco/claude-skills*

    Voice

    Helped on Gemini 3.1 Flash Lite; no reliable difference on Claude Haiku 4.5 and GPT-5 mini. Measured against no skill loaded.

    A judge preferred its output 21 times in 30 over the no-skill version. Instruction arm: not applicable to a blind preference.

    • Claude Haiku 4.5 · unclear
    • Gemini 3.1 Flash Lite · helped
    • GPT-5 mini · unclear

    Cost per task, with the skill loaded: about $0.00334. Adds about 5,700 tokens of context to every request.

The numbers

How to read these tables

How to read these tables

Without -> with
The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
Effect (range)
With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
Orbit
  • Stable: delta at or above the pass threshold, and the interval excludes zero
  • Past the horizon: delta at or below the failure floor, and the interval excludes zero
  • In free drift: the interval spans zero, or the delta sits between the floor and the threshold
  • Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound

A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.

Cohort
The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
Commit
The exact version tested. Later commits are not covered by this verdict.
Injected context tokens
How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
Tested
The date. Verdicts age as models change.

Voice

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded

Via API tested via API

skills/brand-voice Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice Gemini 3.1 Flash Lite First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice GPT-5 mini First run, with and without the skill, Blind comparison of the same task

Holds

No skill: 0.200. With the skill: 0.800. Change +0.600. It could sit between +0.077 and +1.123.

skills/brand-voice Claude Haiku 4.5 First run, with and without the skill, Blind comparison of the same task

No measured effect

No skill: 0.300. With the skill: 0.700. Change +0.400. It could sit between -0.199 and +0.999.

skills/brand-voice GPT-5 mini First run, with and without the skill, Blind comparison of the same task

No measured effect

No skill: 0.400. With the skill: 0.600. Change +0.200. It could sit between -0.440 and +0.840.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.
SkillCohortWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loadedbaseline25%0 by definition02026-08-30
One-line instruction insteadbaselineinstruction arm: not applicable to a blind preference0 by definition02026-08-30
skills/brand-voice from affaan-m/ECC d8409a4b0813 Best result in this testaffaan-m/ECC20% -> 80% baseline shown as the complement+0.60 [+0.60, +0.60]Stable1,1702026-08-30
skills/brand-voice from rampstackco/claude-skills 047924252254rampstackco/claude-skills*30% -> 70% baseline shown as the complement+0.40 [+0.20, +0.60]In free drift on 2 of 35,6742026-08-30
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.