AI Model Comparisons, Tested on the Same Tasks
Each page sets models side by side only where they did exactly the same work. Scores from two test methods are never added together.
| Comparison | What it found |
|---|---|
| Claude Code vs Codex: The Same Tests in Both | The same tips through both. No model was measured in both, so they are described side by side and not ranked. |
| Claude Sonnet vs Opus, Tested on the Same Tasks | 11 jobs with the same work, tested inside Claude Code. They could not be told apart on 5 of the 7 that could be ranked. |
| Claude Opus vs Sonnet vs Haiku, Tested | Eight jobs with the same work, tested inside Claude Code. They could not be told apart on 3 of the 7 that could be ranked. |
| Claude vs GPT, Tested on the Same Tasks | Five jobs with the same work, tested through the API. They could not be told apart on 3 of the 5 that could be ranked. |
Why some pairs have no page
A pair of models gets a page only when they did the same tasks, the same way, on at least three jobs. Fewer than that is not enough to say how they differ. Every model on every job is on the models page, and the same lines are in a data file.