Model under test
Claude Sonnet 5
Every result this site holds for Claude Sonnet 5, reported per instrument. It was measured tested in Claude Code, and the two are never averaged together.
Version string recorded, as returned: claude-sonnet-5. The records carry the requested name back, with no dated snapshot behind it.
What was measured, and where
20 cells in total, and they are counted per instrument below rather than pooled. A cell measured tested via API and a cell measured tested in Claude Code are two measurements, so a single figure across them would be one neither run produced. How the two instruments were compared.
Prompting claims, replayed in Claude Code
12 cells, tested in Claude Code, single-turn replay.
- Unobservable 5 of 12
- In free drift 4 of 12
- Stable 3 of 12
Skills, in Claude Code: Panel v2 (repeat sampling)
8 cells, tested in Claude Code, panel v2 (repeat sampling).
- In free drift 8 of 8
Thinking was refused, and the refusal held on every call
This model was registered to run with thinking suppressed, and a nonzero thinking-token count on it was declared a gate failure that stops the run. Across all of its records the count is zero throughout. That is a gate holding rather than a measured zero, and it is not comparable with a model that was allowed to think.
- Records reporting a nonzero count
- 0 of 1330
- Thinking tokens recorded
- 0
How this model was called, tested in Claude Code
On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this instrument returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.
These cells are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.