Does Brand voice (affaan-m) help? Tested on voice tasks
Brand voice (affaan-m), from affaan-m/ECC. Its voice tasks were run with the skill loaded and without it. Each time it is the same task, run twice.
What we tested
Whether loading this skill helps a model write in a brand's voice.
What counts as helping
A blind reader has to prefer the version written with the skill at least 60 times in 100, across the same 10 briefs.
How we scored it
A blind reader compared the two versions and picked one, without being told which was which. The score is how often the version written with the skill was the one picked, from 0 to 1.
Without to with, per model
Via API tested via API
Claude Haiku 4.5 Blind comparison of the same task
Holds
Gemini 3.1 Flash Lite Blind comparison of the same task
Holds
GPT-5 mini Blind comparison of the same task
Holds
deterministic pass rate, 0 to 1
Show per-task detail
Every item, without and with, per model
Via API tested via API
- Claude Haiku 4.5 Blind comparison of the same task
- Gemini 3.1 Flash Lite Blind comparison of the same task
- GPT-5 mini Blind comparison of the same task
deterministic pass rate, 0 to 1
Hover or focus an item to read its task and every model’s two values.
Task-paired: these item means are what this cell's interval is built from.
The item means behind this plot
| Item | Claude Haiku 4.5 Blind comparison of the same task | Gemini 3.1 Flash Lite Blind comparison of the same task | GPT-5 mini Blind comparison of the same task | |||
|---|---|---|---|---|---|---|
| Task | without | with | without | with | without | with |
| vo-01 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-02 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-03 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-05 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-06 | 0.000 | 1.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-04 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| vo-07 | 0.000 | 1.000 | 1.000 | 0.000 | 0.000 | 1.000 |
| vo-08 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| vo-10 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| vo-09 | 1.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 |
What exactly was tested, and how it was scored
- Repository
- affaan-m/ECC
- Path
- skills/brand-voice
- Commit
d8409a4b0813771235555e32e3d8046a73988bfa- Content hash
eae455eed766cf8ab9d2859713386c56fdce39a8e6a80b60a006a0a076e1d35a- Date tested
- 2026-08-30
The pin is the whole of this skill's identity here. It resolves at https://github.com/affaan-m/ECC/tree/d8409a4b0813771235555e32e3d8046a73988bfa/skills/brand-voice, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.
S10-ecc-brand-voice
Loading skills/brand-voice from affaan-m/ECC improves outputs on voice tasks.
- Pass criterion
- With-arm win rate at or above 0.60 across the same 10 briefs, with the interval excluding a rate of 0.50.
- Scale
- unit. Same instrument and scale as S09.
- Task pairs planned per model
- 10
- Notes
- Shares the voice task set with S09. Paired blind comparison only, and labelled as such on every surface that renders it.
Injected context tokens
| Arm | Context characters | Injected context tokens |
|---|---|---|
| without | 0 | 0 |
| with | 4,780 | 1,122 |
The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.
Verdict per model
| Model | Reading | Orbit | Without -> with | Effect | 95 percent interval | Task pairs | Model version returned |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | Blind comparison of the same task, run twice, with ties counting half | Stable | 0.2000 -> 0.8000 baseline shown as the complement | win rate 0.8000, +0.6000 on the arm difference | [0.0773, 1.1227] | 10 pairs, 20 of 20 records | claude-haiku-4-5-20251001 |
| Gemini 3.1 Flash Lite | Blind comparison of the same task, run twice, with ties counting half | Stable | 0.2000 -> 0.8000 baseline shown as the complement | win rate 0.8000, +0.6000 on the arm difference | [0.0773, 1.1227] | 10 pairs, 20 of 20 records | gemini-3.1-flash-lite |
| GPT-5 mini | Blind comparison of the same task, run twice, with ties counting half | Stable | 0.2000 -> 0.8000 baseline shown as the complement | win rate 0.8000, +0.6000 on the arm difference | [0.0773, 1.1227] | 10 pairs, 20 of 20 records | gpt-5-mini-2025-08-07 |
- Stable: the treatment arm outscored the control arm by more than the threshold, and the interval excludes zero.
Cost
| Model | Arm | Tasks attempted | Mean input tokens | Mean output tokens | Cost per task | Cost ratio | Grading cost per task pair |
|---|---|---|---|---|---|---|---|
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | without | 10 | 70 | 157 | $0.00086 | ||
Claude Haiku 4.5 served claude-haiku-4-5-20251001 | with | 10 | 1,250 | 151 | $0.00200 | 2.34x | $0.00223 |
GPT-5 mini served gpt-5-mini-2025-08-07 | without | 10 | 64 | 142 | $0.00030 | ||
GPT-5 mini served gpt-5-mini-2025-08-07 | with | 10 | 1,109 | 131 | $0.00054 | 1.79x | $0.00213 |
| Gemini 3.1 Flash Lite | without | 10 | 63 | 138 | $0.00022 | ||
| Gemini 3.1 Flash Lite | with | 10 | 1,153 | 141 | $0.00050 | 2.25x | $0.00222 |
Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.
Where the with arm did not win
Claude Haiku 4.5, Blind comparison of the same task, run twice, with ties counting half
Lost on 2 of 10 pairs: vo-09 (-1.0000), vo-10 (-1.0000).
Gemini 3.1 Flash Lite, Blind comparison of the same task, run twice, with ties counting half
Lost on 2 of 10 pairs: vo-07 (-1.0000), vo-09 (-1.0000).
GPT-5 mini, Blind comparison of the same task, run twice, with ties counting half
Lost on 2 of 10 pairs: vo-04 (-1.0000), vo-08 (-1.0000).
Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.