Model under test
Claude Fable 5
Every result this site holds for Claude Fable 5, reported per instrument. It was measured tested in Claude Code, and the two are never averaged together.
Version strings recorded, as returned: claude-fable-5, claude-opus-5.
What was measured, and where
26 cells in total, and they are counted per instrument below rather than pooled. A cell measured tested via API and a cell measured tested in Claude Code are two measurements, so a single figure across them would be one neither run produced. How the two instruments were compared.
Prompting claims, replayed in Claude Code
12 cells, tested in Claude Code, single-turn replay.
- Unobservable 5 of 12
- In free drift 4 of 12
- Stable 3 of 12
Skills, in Claude Code: Skill cohort (one pass)
6 cells, tested in Claude Code, skill cohort (one pass).
- In free drift 5 of 6
- Unobservable 1 of 6
Skills, in Claude Code: Panel v2 (repeat sampling)
8 cells, tested in Claude Code, panel v2 (repeat sampling).
- In free drift 7 of 8
- Unobservable 1 of 8
Thinking was recorded rather than refused, by prior registration
Alone among the models on this instrument, this one was registered in advance to have its thinking RECORDED rather than suppressed, and a nonzero count on it is not a gate failure. That was registered before the run precisely so that a result on this model could not later be explained away, or quietly discarded, on the grounds that it thought. The figures beside this are a measurement; the zeroes on the other models are a refusal holding, and the two are not the same kind of number.
- Records reporting a nonzero count
- 1048 of 1358
- Thinking tokens recorded
- 148,834
The plan served a different model on some calls, and those calls were refused
Some calls asking for this model came back answered by another one: same prompts, same flags, no error, and a plausible score. Nothing in the envelope announced it except the served model string the adapter now records on every record. A record served by another model is evidence about that other model, gathered under conditions chosen for this one, so it is invalid and never scored. It is not deleted and its tokens still count against the run's spend, because the call was made whoever answered it.
- Calls served by another model
- 48 of 1358 on the tips replay, spanning 28 unit-runs across 10 distinct claim, task and arm unitsServed instead by Claude Opus 5, on 2026-08-31.
- Skill cells withheld rather than thinned
- 2Each is published as withheld, with its reason and its surviving pair count, on the skill’s own page.
- Prevented from
- 2026-08-31The served model is read back from the CLI’s own usage envelope and recorded on every record, and a record whose served model differs from the requested one is invalid at birth and invalid at load. Committed in
7ac6030and81253fc.
How this model was called, tested in Claude Code
On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this instrument returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.
These cells are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.