AI Model Comparisons, Tested on the Same Tasks

Each page sets models side by side only where they did exactly the same work. Scores from two test methods are never added together.

ComparisonWhat it found
Claude Code vs Codex: The Same Tests in BothThe same tips through both. No model was measured in both, so they are described side by side and not ranked.
Claude Sonnet vs Opus, Tested on the Same Tasks11 jobs with the same work, tested inside Claude Code. They could not be told apart on 5 of the 7 that could be ranked.
Claude Opus vs Sonnet vs Haiku, TestedEight jobs with the same work, tested inside Claude Code. They could not be told apart on 3 of the 7 that could be ranked.
Claude vs GPT, Tested on the Same TasksFive jobs with the same work, tested through the API. They could not be told apart on 3 of the 5 that could be ranked.
Why some pairs have no page

A pair of models gets a page only when they did the same tasks, the same way, on at least three jobs. Fewer than that is not enough to say how they differ. Every model on every job is on the models page, and the same lines are in a data file.