Best LLM for Step-by-Step Answers, Tested
Step-by-step answers means a problem with one exact answer that takes several steps to reach. Each model was given the same ten tasks with no help, tested in the API, Claude Code and Codex.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
Gemini 3.1 Flash Lite 40.0%
Too close to call with Claude Haiku 4.5
tier 1, Main run, tip by tip In Claude Code
Claude Fable 5 40.0%, Claude Fable 5.1 40.0%, Claude Opus 5 40.0%, Claude Sonnet 5 40.0%, Claude Opus 5.5 40.0%, tied
Too close to call with Claude Haiku 4.5
tier 1, Main run
Some of these models come from another run of the same tasks In Codex
GPT-6 Sol 40.0%
No range published, so nobody is placed against it
tier 1, Codex run, tip by tip | Via API
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. In Claude Code
Claude Haiku 4.5 $0.00135 a task
Against Claude Sonnet 5 $0.00630 a task, the best score
run on subscription; shown at API list price for comparison In Codex
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked. In Claude Code
Claude Fable 5, same answer on 10 of 10; Claude Fable 5.1, same answer on 10 of 10; Claude Opus 5, same answer on 10 of 10; Claude Opus 5.5, same answer on 10 of 10; Claude Sonnet 5, same answer on 10 of 10, tied
Too close to call with Claude Haiku 4.5
no sampling control here, Each tip asked five times In Codex
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
Gemini 3.1 Flash Lite 563 ms
Median time per call, across all twelve tips asked five times
2 models timed, Each tip asked five times In Claude Code
Claude Haiku 4.5 11.4 s
Median time per call, across all twelve tips asked five times
6 models timed, Each tip asked five times In Codex
No run of this job timed its calls this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 30.0% tier 1 too close to call | 20 of 20 at temperature 0 |
| GPT-5 mini | 16.0% tier 1 | 10 of 10 at the vendor default |
| Gemini 3.1 Flash Lite | 40.0% tier 1 best | 10 of 10 at temperature 0 |
| GPT-5.4 mini | not run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 20.0% tier 1 too close to call | 8 of 10 no sampling control here |
| Claude Fable 5 | 40.0% tier 1 best | 10 of 10 no sampling control here |
| Claude Fable 5.1 | 40.0% tier 1 best | 10 of 10 no sampling control here |
| Claude Opus 5 | 40.0% tier 1 best | 10 of 10 no sampling control here |
| Claude Opus 5.5 | 40.0% tier 1 best | 10 of 10 no sampling control here |
| Claude Sonnet 5 | 40.0% tier 1 best | 10 of 10 no sampling control here |
In Codex
| Model | Score | Same answer twice |
|---|---|---|
| GPT-5.4 mini | 30.0% tier 1 | one pass, not measured |
| GPT-5.6 Luna | 30.0% tier 1 | one pass, not measured |
| GPT-5.6 Terra | 30.0% tier 1 | one pass, not measured |
| GPT-6 Luna | 30.0% tier 1 | one pass, not measured |
| GPT-6 Sol | 40.0% tier 1 best | one pass, not measured |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Better step-by-step answers | Claude Haiku 4.5 | +20.0 percentage points (+1.0 to +39.0) | no one-line instruction was tested on tips |
| Better step-by-step answers | GPT-5 mini | +34.0 percentage points (+16.6 to +51.4) | no one-line instruction was tested on tips |
In Claude Code
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Better step-by-step answers | Claude Haiku 4.5 | +30.0 percentage points (+12.1 to +47.9) | no one-line instruction was tested on tips |
How hard the tasks were
These scores are from the first tasks we built for this job, tier 1, the set the models are read on.
The full tables behind this page
Via API
Read from Main run, tip by tip, 2026-08-17.
In Claude Code
Read from Main run, 2026-09-03, with the same tasks also from Claude Opus 5.5, asked each task five times, 2026-09-28.
- Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 40.0% (tier 1)
In Codex
Read from Codex run, tip by tip, 2026-10-03.
| Tip | Model | Test method | Run |
|---|---|---|---|
| Better step-by-step answers | Claude Haiku 4.5 | Via API | Main run, tier 1 |
| Better step-by-step answers | GPT-5 mini | Via API | Main run, tier 1 |
| Better step-by-step answers | Claude Haiku 4.5 | In Claude Code | Main run, tier 1 |
Version pairs on this job
- On better step-by-step answers, Claude Fable 5.1 scores 40% unaided against 40% for Claude Fable 5; no measurable change.
The job in full: Better step-by-step answers. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/step-by-step.json.