Model under test

Claude Fable 5

Every result this site holds for Claude Fable 5, reported per instrument. It was measured tested in Claude Code, and the two are never averaged together.

Version strings recorded, as returned: claude-fable-5, claude-opus-5.

What was measured, and where

26 cells in total, and they are counted per instrument below rather than pooled. A cell measured tested via API and a cell measured tested in Claude Code are two measurements, so a single figure across them would be one neither run produced. How the two instruments were compared.

Prompting claims, replayed in Claude Code

12 cells, tested in Claude Code, single-turn replay.

  • Unobservable 5 of 12
  • In free drift 4 of 12
  • Stable 3 of 12

Skills, in Claude Code: Skill cohort (one pass)

6 cells, tested in Claude Code, skill cohort (one pass).

  • In free drift 5 of 6
  • Unobservable 1 of 6

Skills, in Claude Code: Panel v2 (repeat sampling)

8 cells, tested in Claude Code, panel v2 (repeat sampling).

  • In free drift 7 of 8
  • Unobservable 1 of 8

Thinking was recorded rather than refused, by prior registration

Alone among the models on this instrument, this one was registered in advance to have its thinking RECORDED rather than suppressed, and a nonzero count on it is not a gate failure. That was registered before the run precisely so that a result on this model could not later be explained away, or quietly discarded, on the grounds that it thought. The figures beside this are a measurement; the zeroes on the other models are a refusal holding, and the two are not the same kind of number.

Records reporting a nonzero count
1048 of 1358
Thinking tokens recorded
148,834

The plan served a different model on some calls, and those calls were refused

Some calls asking for this model came back answered by another one: same prompts, same flags, no error, and a plausible score. Nothing in the envelope announced it except the served model string the adapter now records on every record. A record served by another model is evidence about that other model, gathered under conditions chosen for this one, so it is invalid and never scored. It is not deleted and its tokens still count against the run's spend, because the call was made whoever answered it.

Calls served by another model
48 of 1358 on the tips replay, spanning 28 unit-runs across 10 distinct claim, task and arm unitsServed instead by Claude Opus 5, on 2026-08-31.
Skill cells withheld rather than thinned
2Each is published as withheld, with its reason and its surviving pair count, on the skill’s own page.
Prevented from
2026-08-31The served model is read back from the CLI’s own usage envelope and recorded on every record, and a record whose served model differs from the requested one is invalid at birth and invalid at load. Committed in 7ac6030 and 81253fc.

The finding in full, and the checks that catch it now.

How this model was called, tested in Claude Code

On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this instrument returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.

These cells are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.

Every figure on this page is computed at build time from the committed records. Nothing here averages a figure from one instrument with a figure from another, and the page carries no spread across them, because a model measured two ways has two results and not one.