Model under test

Claude Sonnet 5

Every result this site holds for Claude Sonnet 5, reported per instrument. It was measured tested in Claude Code, and the two are never averaged together.

Version string recorded, as returned: claude-sonnet-5. The records carry the requested name back, with no dated snapshot behind it.

What was measured, and where

20 cells in total, and they are counted per instrument below rather than pooled. A cell measured tested via API and a cell measured tested in Claude Code are two measurements, so a single figure across them would be one neither run produced. How the two instruments were compared.

Prompting claims, replayed in Claude Code

12 cells, tested in Claude Code, single-turn replay.

  • Unobservable 5 of 12
  • In free drift 4 of 12
  • Stable 3 of 12

Skills, in Claude Code: Panel v2 (repeat sampling)

8 cells, tested in Claude Code, panel v2 (repeat sampling).

  • In free drift 8 of 8

Thinking was refused, and the refusal held on every call

This model was registered to run with thinking suppressed, and a nonzero thinking-token count on it was declared a gate failure that stops the run. Across all of its records the count is zero throughout. That is a gate holding rather than a measured zero, and it is not comparable with a model that was allowed to think.

Records reporting a nonzero count
0 of 1330
Thinking tokens recorded
0

How this model was called, tested in Claude Code

On the Claude Code CLI, on a subscription path with no API key and nothing billed. Every cost figure this instrument returns is a client-side estimate at list price and is recorded under that label and no other. What that means, and how it was calibrated.

These cells are tested in Claude Code and are never averaged with any figure tested via API. Where this model was measured on both, the two are reported as two panels above.

Every figure on this page is computed at build time from the committed records. Nothing here averages a figure from one instrument with a figure from another, and the page carries no spread across them, because a model measured two ways has two results and not one.