Does Brand voice (rampstackco) help? Tested on voice tasks
Brand voice (rampstackco), from rampstackco/claude-skills. Its voice tasks were run with the skill loaded and without it. Each time it is the same task, run twice.
What we tested
Whether loading this skill helps a model write in a brand's voice.
What counts as helping
A blind reader has to prefer the version written with the skill at least 60 times in 100, across the same 10 briefs.
How we scored it
A blind reader compared the two versions and picked one, without being told which was which. The score is how often the version written with the skill was the one picked, from 0 to 1.
Without to with, per model
Via API tested via API
Gemini 3.1 Flash Lite Blind comparison of the same task
Holds
Claude Haiku 4.5 Blind comparison of the same task
No measured effect
GPT-5 mini Blind comparison of the same task
No measured effect
deterministic pass rate, 0 to 1
Show per-task detail
Every item, without and with, per model
Via API tested via API
- Gemini 3.1 Flash Lite Blind comparison of the same task
- Claude Haiku 4.5 Blind comparison of the same task
- GPT-5 mini Blind comparison of the same task
deterministic pass rate, 0 to 1
Hover or focus an item to read its task and every model’s two values.
Task-paired: these item means are what this cell's interval is built from.
The item means behind this plot
| Item | Gemini 3.1 Flash Lite Blind comparison of the same task | Claude Haiku 4.5 Blind comparison of the same task | GPT-5 mini Blind comparison of the same task | |||
|---|---|---|---|---|---|---|
| Task | without | with | without | with | without | with |
| vo-01 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-02 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-04 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-06 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-10 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-03 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| vo-08 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| vo-05 | 0.000 | 1.000 | 1.000 | 0.000 | 1.000 | 0.000 |
| vo-09 | 1.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 |
| vo-07 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 |
What exactly was tested, and how it was scored
- Repository
- rampstackco/claude-skills
- Path
- skills/brand-voice
- Commit
0479242522549dfdb389bb9b7807ad4d6016ffb7- Content hash
05eb2402e3e2459be9809c698160bb83702bfcb53c871df1cd9e18f59fb5c0cd- Date tested
- 2026-08-30
The pin is the whole of this skill's identity here. It resolves at https://github.com/rampstackco/claude-skills/tree/0479242522549dfdb389bb9b7807ad4d6016ffb7/skills/brand-voice, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.
S09-brand-voice
Loading skills/brand-voice from rampstackco/claude-skills improves outputs on voice tasks.
- Pass criterion
- With-arm win rate at or above 0.60 across the same 10 briefs, with the interval excluding a rate of 0.50.
- Scale
- unit. A win rate, bounded at 0 and 1, which is the unit scale. The bridge from a preference to the two arm means status_v1 needs is stated once in src/scoring/paired.ts and nowhere else.
- Task pairs planned per model
- 10
- Notes
- PAIRED BLIND COMPARISON ONLY. There is no deterministic score for voice and none is invented. The grader is never asked how good an output is, only which of two it prefers, so no absolute score exists to render.
Injected context tokens
| Arm | Context characters | Injected context tokens |
|---|---|---|
| without | 0 | 0 |
| with | 22,306 | 5,236 |
The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.
Verdict per model
| Model | Reading | Orbit | Without -> with | Effect | 95 percent interval | Task pairs | Model version returned |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | Blind comparison of the same task, run twice, with ties counting half | In free drift unclear. The model scored 30% without it. | 0.3000 -> 0.7000 baseline shown as the complement | win rate 0.7000, +0.4000 on the arm difference | [-0.1988, 0.9988] | 10 pairs, 20 of 20 records | claude-haiku-4-5-20251001 |
| Gemini 3.1 Flash Lite | Blind comparison of the same task, run twice, with ties counting half | Stable | 0.2000 -> 0.8000 baseline shown as the complement | win rate 0.8000, +0.6000 on the arm difference | [0.0773, 1.1227] | 10 pairs, 20 of 20 records | gemini-3.1-flash-lite |
| GPT-5 mini | Blind comparison of the same task, run twice, with ties counting half | In free drift unclear. The model scored 40% without it. | 0.4000 -> 0.6000 baseline shown as the complement | win rate 0.6000, +0.2000 on the arm difference | [-0.4401, 0.8401] | 10 pairs, 20 of 20 records | gpt-5-mini-2025-08-07 |
- In free drift: no separation the design can resolve. Not evidence of no effect.
- Stable: the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero.
Cost
| Model | Arm | Tasks attempted | Mean input tokens | Mean output tokens | Cost per task | Cost ratio | Grading cost per task pair |
|---|---|---|---|---|---|---|---|
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | without | 10 | 70 | 156 | $0.00085 | ||
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | with | 10 | 6,046 | 148 | $0.00679 | 7.99x | $0.00220 |
GPT-5 mini served gpt-5-mini-2025-08-07 | without | 10 | 64 | 137 | $0.00029 | ||
GPT-5 mini served gpt-5-mini-2025-08-07 | with | 10 | 5,328 | 141 | $0.00161 | 5.56x | $0.00213 |
| Gemini 3.1 Flash Lite | without | 10 | 63 | 138 | $0.00022 | ||
| Gemini 3.1 Flash Lite | with | 10 | 5,649 | 136 | $0.00162 | 7.27x | $0.00219 |
Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.
Where the with arm did not win
Claude Haiku 4.5, Blind comparison of the same task, run twice, with ties counting half
Lost on 3 of 10 pairs: vo-05 (-1.0000), vo-07 (-1.0000), vo-09 (-1.0000).
Gemini 3.1 Flash Lite, Blind comparison of the same task, run twice, with ties counting half
Lost on 2 of 10 pairs: vo-07 (-1.0000), vo-09 (-1.0000).
GPT-5 mini, Blind comparison of the same task, run twice, with ties counting half
Lost on 4 of 10 pairs: vo-03 (-1.0000), vo-05 (-1.0000), vo-07 (-1.0000), vo-08 (-1.0000).
Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.