Which LLM Hits an Exact Word Count? Tested

Getting the length right means asking for an answer of an exact number of words and getting it. Each model was given the same ten tasks with no help, tested in the API, Claude Code and Codex.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API GPT-5 mini 5.0% Too close to call with Gemini 3.1 Flash Lite tier 2, Harder tasks, tier 2
In Claude Code Claude Haiku 4.5 4.2% Too close to call with Claude Opus 5, Claude Sonnet 5, Claude Fable 5.1 and Claude Opus 5.5 tier 2, Harder tasks, tier 2 Some of these models come from another run of the same tasks
In Codex GPT-5.6 Terra 62.4% No range published, so nobody is placed against it tier 1, Codex run, tip by tip
Via API This run does not publish a cost per task.
In Claude Code This run does not publish a cost per task.
In Codex The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
Via API The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked.
In Claude Code Claude Opus 5, same answer on 8 of 10 Too close to call with Claude Haiku 4.5, Claude Fable 5.1 and Claude Sonnet 5 no sampling control here, Each tip asked five times
In Codex No run asked these tasks more than once this way, so there is no repeat to compare.
Via API Gemini 3.1 Flash Lite 563 ms Median time per call, across all twelve tips asked five times 2 models timed, Each tip asked five times
In Claude Code Claude Haiku 4.5 11.4 s Median time per call, across all twelve tips asked five times 6 models timed, Each tip asked five times
In Codex No run of this job timed its calls this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.5in another run
GPT-5 mini5.0% tier 2 best8 of 10 at the vendor default
Gemini 3.1 Flash Lite3.7% tier 2 too close to call10 of 10 at temperature 0
GPT-5.4 miniin another run

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.54.2% tier 2 best7 of 10 no sampling control here
Claude Fable 5in another run
Claude Fable 5.10.3% tier 2 too close to call4 of 10 no sampling control here
Claude Opus 50.3% tier 2 too close to call8 of 10 no sampling control here
Claude Opus 5.50.3% tier 2 too close to call0 of 10 no sampling control here
Claude Sonnet 51.4% tier 2 too close to call7 of 10 no sampling control here

In Codex

ModelScoreSame answer twice
GPT-5.4 mini20.6% tier 1one pass, not measured
GPT-5.6 Luna57.7% tier 1one pass, not measured
GPT-5.6 Terra62.4% tier 1 bestone pass, not measured
GPT-6 Luna51.0% tier 1one pass, not measured
GPT-6 Solwithheld: Incomplete coverage at retry exhaustion; 1 registered units have no valid record. Resume only within registered retry policy.

What measurably helps on this job

Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.

Via API

TipModelBetter than no help byAgainst one plain sentence
Getting the length rightGPT-5 mini+86.5 percentage points (+76.7 to +96.3)no one-line instruction was tested on tips
Getting the length rightGemini 3.1 Flash Lite+95.4 percentage points (+88.7 to +102.1)no one-line instruction was tested on tips

In Claude Code

TipModelBetter than no help byAgainst one plain sentence
Getting the length rightClaude Opus 5+93.1 percentage points (+90.2 to +95.9)no one-line instruction was tested on tips
Getting the length rightClaude Sonnet 5+93.3 percentage points (+88.2 to +98.5)no one-line instruction was tested on tips
Getting the length rightClaude Haiku 4.5+77.4 percentage points (+67.9 to +86.9)no one-line instruction was tested on tips
Getting the length rightClaude Fable 5.1+99.8 percentage points (+99.3 to +100.2)no one-line instruction was tested on tips
Getting the length rightClaude Opus 5.5+99.8 percentage points (+99.3 to +100.2)no one-line instruction was tested on tips

In Codex

TipModelBetter than no help byAgainst one plain sentence
Getting the length rightGPT-5.6 Terra+66.0 percentage points (+45.1 to +86.9)no one-line instruction was tested on tips
Getting the length rightGPT-5.6 Luna+62.6 percentage points (+38.3 to +86.8)no one-line instruction was tested on tips
Getting the length rightGPT-6 Sol+71.7 percentage points (+51.2 to +92.1)no one-line instruction was tested on tips
Getting the length rightGPT-6 Luna+79.3 percentage points (+68.2 to +90.3)no one-line instruction was tested on tips

How hard the tasks were

The first tasks got too easy to tell the models apart, so we built a harder set, tier 2, and the scores here are from it. In Codex the harder set was not run, so those scores are from tier 1.

The full tables behind this page

Via API

Read from Harder tasks, tier 2, 2026-09-19.

  • Another run, Main run, tip by tip, 2026-08-17: Gemini 3.1 Flash Lite 31.0% (tier 1), Claude Haiku 4.5 24.7% (tier 1), GPT-5 mini 1.6% (tier 1)
  • Another run, API twin for the Codex run, 2026-09-07: GPT-5.4 mini 15.7% (tier 1)

In Claude Code

Read from Harder tasks, tier 2, 2026-09-19, with the same tasks also from Claude Opus 5.5, tier 2, 2026-09-25.

  • Another run, Main run, 2026-09-03: Claude Fable 5 76.0% (tier 1), Claude Fable 5.1 50.2% (tier 1), Claude Haiku 4.5 27.5% (tier 1), Claude Sonnet 5 17.8% (tier 1), Claude Opus 5 15.0% (tier 1)
  • Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 80.1% (tier 1)
  • Another run, Claude Opus 5.5, asked each task five times, 2026-09-28: Claude Opus 5.5 76.6% (tier 1)

In Codex

Read from Codex run, tip by tip, 2026-10-03.

  • Another run, Codex run, replaying two earlier tests, 2026-09-07: GPT-5.4 mini 18.6% (tier 1)
Every result above, with the run it was read from.
TipModelTest methodRun
Getting the length rightClaude Opus 5In Claude Codetier 2, tier 2
Getting the length rightClaude Sonnet 5In Claude Codetier 2, tier 2
Getting the length rightClaude Haiku 4.5In Claude Codetier 2, tier 2
Getting the length rightGPT-5 miniVia APItier 2, tier 2
Getting the length rightGemini 3.1 Flash LiteVia APItier 2, tier 2
Getting the length rightClaude Fable 5.1In Claude Codetier 2, tier 2
Getting the length rightGPT-5.6 TerraIn Codextier 2, tier 2
Getting the length rightGPT-5.6 LunaIn Codextier 2, tier 2
Getting the length rightGPT-6 SolIn Codextier 2, tier 2
Getting the length rightGPT-6 LunaIn Codextier 2, tier 2
Getting the length rightClaude Opus 5.5In Claude CodeClaude Opus 5.5, tier 2, tier 2

Version pairs on this job

  • On getting the length right, Claude Fable 5.1 scores 50% unaided against 76% for Claude Fable 5; scores lower.

The job in full: Getting the length right. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/length-control.json.