Best LLM for Step-by-Step Answers, Tested

Step-by-step answers means a problem with one exact answer that takes several steps to reach. Each model was given the same ten tasks with no help, tested in the API, Claude Code and Codex.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API Gemini 3.1 Flash Lite 40.0% Too close to call with Claude Haiku 4.5 tier 1, Main run, tip by tip
In Claude Code Claude Fable 5 40.0%, Claude Fable 5.1 40.0%, Claude Opus 5 40.0%, Claude Sonnet 5 40.0%, Claude Opus 5.5 40.0%, tied Too close to call with Claude Haiku 4.5 tier 1, Main run Some of these models come from another run of the same tasks
In Codex GPT-6 Sol 40.0% No range published, so nobody is placed against it tier 1, Codex run, tip by tip
Via API The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
In Claude Code Claude Haiku 4.5 $0.00135 a task Against Claude Sonnet 5 $0.00630 a task, the best score run on subscription; shown at API list price for comparison
In Codex The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked.
In Claude Code Claude Fable 5, same answer on 10 of 10; Claude Fable 5.1, same answer on 10 of 10; Claude Opus 5, same answer on 10 of 10; Claude Opus 5.5, same answer on 10 of 10; Claude Sonnet 5, same answer on 10 of 10, tied Too close to call with Claude Haiku 4.5 no sampling control here, Each tip asked five times
In Codex No run asked these tasks more than once this way, so there is no repeat to compare.
Via API Gemini 3.1 Flash Lite 563 ms Median time per call, across all twelve tips asked five times 2 models timed, Each tip asked five times
In Claude Code Claude Haiku 4.5 11.4 s Median time per call, across all twelve tips asked five times 6 models timed, Each tip asked five times
In Codex No run of this job timed its calls this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.530.0% tier 1 too close to call20 of 20 at temperature 0
GPT-5 mini16.0% tier 110 of 10 at the vendor default
Gemini 3.1 Flash Lite40.0% tier 1 best10 of 10 at temperature 0
GPT-5.4 mininot run

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.520.0% tier 1 too close to call8 of 10 no sampling control here
Claude Fable 540.0% tier 1 best10 of 10 no sampling control here
Claude Fable 5.140.0% tier 1 best10 of 10 no sampling control here
Claude Opus 540.0% tier 1 best10 of 10 no sampling control here
Claude Opus 5.540.0% tier 1 best10 of 10 no sampling control here
Claude Sonnet 540.0% tier 1 best10 of 10 no sampling control here

In Codex

ModelScoreSame answer twice
GPT-5.4 mini30.0% tier 1one pass, not measured
GPT-5.6 Luna30.0% tier 1one pass, not measured
GPT-5.6 Terra30.0% tier 1one pass, not measured
GPT-6 Luna30.0% tier 1one pass, not measured
GPT-6 Sol40.0% tier 1 bestone pass, not measured

What measurably helps on this job

Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.

Via API

TipModelBetter than no help byAgainst one plain sentence
Better step-by-step answersClaude Haiku 4.5+20.0 percentage points (+1.0 to +39.0)no one-line instruction was tested on tips
Better step-by-step answersGPT-5 mini+34.0 percentage points (+16.6 to +51.4)no one-line instruction was tested on tips

In Claude Code

TipModelBetter than no help byAgainst one plain sentence
Better step-by-step answersClaude Haiku 4.5+30.0 percentage points (+12.1 to +47.9)no one-line instruction was tested on tips

How hard the tasks were

These scores are from the first tasks we built for this job, tier 1, the set the models are read on.

The full tables behind this page

Via API

Read from Main run, tip by tip, 2026-08-17.

In Claude Code

Read from Main run, 2026-09-03, with the same tasks also from Claude Opus 5.5, asked each task five times, 2026-09-28.

  • Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 40.0% (tier 1)

In Codex

Read from Codex run, tip by tip, 2026-10-03.

Every result above, with the run it was read from.
TipModelTest methodRun
Better step-by-step answersClaude Haiku 4.5Via APIMain run, tier 1
Better step-by-step answersGPT-5 miniVia APIMain run, tier 1
Better step-by-step answersClaude Haiku 4.5In Claude CodeMain run, tier 1

Version pairs on this job

  • On better step-by-step answers, Claude Fable 5.1 scores 40% unaided against 40% for Claude Fable 5; no measurable change.

The job in full: Better step-by-step answers. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/step-by-step.json.