Does Brand voice (rampstackco) help? Tested on voice tasks

Brand voice (rampstackco), from rampstackco/claude-skills. Its voice tasks were run with the skill loaded and without it. Each time it is the same task, run twice.

What we tested

Whether loading this skill helps a model write in a brand's voice.

What counts as helping

A blind reader has to prefer the version written with the skill at least 60 times in 100, across the same 10 briefs.

How we scored it

A blind reader compared the two versions and picked one, without being told which was which. The score is how often the version written with the skill was the one picked, from 0 to 1.

Without to with, per model

Via API tested via API

Gemini 3.1 Flash Lite Blind comparison of the same task

0.200
0.800

Holds

Claude Haiku 4.5 Blind comparison of the same task

0.300
0.700

No measured effect

GPT-5 mini Blind comparison of the same task

0.400
0.600

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Grey is the score without. Colour is the score with. The bracket shows how much the difference could move if we ran it again. Rows are ordered by the score with. A model measured at two versions keeps its versions next to each other. The two test methods are reported separately and never averaged. A bracket wider than the axis is drawn to the edge with its cap omitted; the table gives its bounds.
Show per-task detail

Every item, without and with, per model

Via API tested via API

  • Gemini 3.1 Flash Lite Blind comparison of the same task
  • Claude Haiku 4.5 Blind comparison of the same task
  • GPT-5 mini Blind comparison of the same task

deterministic pass rate, 0 to 1

Hover or focus an item to read its task and every model’s two values.

Task-paired: these item means are what this cell's interval is built from.

The item means behind this plot
Per-item arm means. Without is the unaided arm, with is the treated arm.
ItemGemini 3.1 Flash Lite Blind comparison of the same task Claude Haiku 4.5 Blind comparison of the same task GPT-5 mini Blind comparison of the same task
Task without with without with without with
vo-01 0.000 1.000 0.000 1.000 0.000 1.000
vo-02 0.000 1.000 0.000 1.000 0.000 1.000
vo-04 0.000 1.000 0.000 1.000 0.000 1.000
vo-06 0.000 1.000 0.000 1.000 0.000 1.000
vo-10 0.000 1.000 0.000 1.000 0.000 1.000
vo-03 0.000 1.000 0.000 1.000 1.000 0.000
vo-08 0.000 1.000 0.000 1.000 1.000 0.000
vo-05 0.000 1.000 1.000 0.000 1.000 0.000
vo-09 1.000 0.000 1.000 0.000 0.000 1.000
vo-07 1.000 0.000 1.000 0.000 1.000 0.000
A picture of the per-item means behind the per-model numbers. The tables are the source. Each model draws two lines over the same task set: a dashed line through its unaided scores and a solid line through its treated ones. Items are ordered by the mean unaided score across the models that measured them, lowest first. The two instruments are reported separately and never averaged.
What exactly was tested, and how it was scored
Repository
rampstackco/claude-skills
Path
skills/brand-voice
Commit
0479242522549dfdb389bb9b7807ad4d6016ffb7
Content hash
05eb2402e3e2459be9809c698160bb83702bfcb53c871df1cd9e18f59fb5c0cd
Date tested
2026-08-30

The pin is the whole of this skill's identity here. It resolves at https://github.com/rampstackco/claude-skills/tree/0479242522549dfdb389bb9b7807ad4d6016ffb7/skills/brand-voice, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.

S09-brand-voice

Loading skills/brand-voice from rampstackco/claude-skills improves outputs on voice tasks.

Pass criterion
With-arm win rate at or above 0.60 across the same 10 briefs, with the interval excluding a rate of 0.50.
Scale
unit. A win rate, bounded at 0 and 1, which is the unit scale. The bridge from a preference to the two arm means status_v1 needs is stated once in src/scoring/paired.ts and nowhere else.
Task pairs planned per model
10
Notes
PAIRED BLIND COMPARISON ONLY. There is no deterministic score for voice and none is invented. The grader is never asked how good an output is, only which of two it prefers, so no absolute score exists to render.

Injected context tokens

Injected context tokens, per arm
ArmContext charactersInjected context tokens
without00
with22,3065,236

The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.

Verdict per model

Effect per model
ModelReadingOrbitWithout -> withEffect95 percent intervalTask pairsModel version returned
Claude Haiku 4.5Blind comparison of the same task, run twice, with ties counting halfIn free drift unclear. The model scored 30% without it.0.3000 -> 0.7000 baseline shown as the complementwin rate 0.7000, +0.4000 on the arm difference[-0.1988, 0.9988]10 pairs, 20 of 20 recordsclaude-haiku-4-5-20251001
Gemini 3.1 Flash LiteBlind comparison of the same task, run twice, with ties counting halfStable0.2000 -> 0.8000 baseline shown as the complementwin rate 0.8000, +0.6000 on the arm difference[0.0773, 1.1227]10 pairs, 20 of 20 recordsgemini-3.1-flash-lite
GPT-5 miniBlind comparison of the same task, run twice, with ties counting halfIn free drift unclear. The model scored 40% without it.0.4000 -> 0.6000 baseline shown as the complementwin rate 0.6000, +0.2000 on the arm difference[-0.4401, 0.8401]10 pairs, 20 of 20 recordsgpt-5-mini-2025-08-07
  • In free drift: no separation the design can resolve. Not evidence of no effect.
  • Stable: the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero.

Cost

Cost per task
ModelArmTasks attemptedMean input tokensMean output tokensCost per taskCost ratioGrading cost per task pair
Claude Haiku 4.5 served claude-haiku-4-5-20251001without1070156$0.00085
Claude Haiku 4.5 served claude-haiku-4-5-20251001with106,046148$0.006797.99x$0.00220
GPT-5 mini served gpt-5-mini-2025-08-07without1064137$0.00029
GPT-5 mini served gpt-5-mini-2025-08-07with105,328141$0.001615.56x$0.00213
Gemini 3.1 Flash Litewithout1063138$0.00022
Gemini 3.1 Flash Litewith105,649136$0.001627.27x$0.00219

Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.

Where the with arm did not win

Claude Haiku 4.5, Blind comparison of the same task, run twice, with ties counting half

Lost on 3 of 10 pairs: vo-05 (-1.0000), vo-07 (-1.0000), vo-09 (-1.0000).

Gemini 3.1 Flash Lite, Blind comparison of the same task, run twice, with ties counting half

Lost on 2 of 10 pairs: vo-07 (-1.0000), vo-09 (-1.0000).

GPT-5 mini, Blind comparison of the same task, run twice, with ties counting half

Lost on 4 of 10 pairs: vo-03 (-1.0000), vo-05 (-1.0000), vo-07 (-1.0000), vo-08 (-1.0000).

Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.

720 run records behind this page. Every verdict is computed at build time by the same frozen status_v1 rule that decides every other verdict on this site, and nothing here is written by hand.