Models measured on this site
Every model that has at least one cell, with the instruments that measured it and the number of task sets it ran. The figures themselves are on the comparison tables, where each is shown within one task set and one instrument and never averaged across them.
Every count on this page is derived from the committed cells. A model appears here because some instrument produced a cell for it, not because it was planned.
| Model | Instruments | Versions | Task sets measured | Page |
|---|---|---|---|---|
| Claude Haiku 4.5 | tested via API; tested in Claude Code | 1 | 19 | Claude Haiku 4.5 |
| GPT-5 mini | tested via API | 1 | 15 | GPT-5 mini |
| Gemini 3.1 Flash Lite | tested via API | 1 | 15 | Gemini 3.1 Flash Lite |
| Claude Fable 5 | tested in Claude Code | 2Claude Fable 5 and Claude Fable 5.1 | 15 | Claude Fable 5 |
| Claude Fable 5.1 | tested in Claude Code | 2Claude Fable 5 and Claude Fable 5.1 | 16 | Claude Fable 5.1 |
| Claude Opus 5 | tested in Claude Code | 1 | 16 | Claude Opus 5 |
| Claude Sonnet 5 | tested in Claude Code | 1 | 16 | Claude Sonnet 5 |
A model line with more than one served version renders as a series: each version keeps its own row on every table, and the per-task-set shifts between them are published on the comparison page and on each version’s own page. Versions are never collapsed into one row.