Which LLM Hits an Exact Word Count? Tested
Getting the length right means asking for an answer of an exact number of words and getting it. Each model was given the same ten tasks with no help, tested in the API, Claude Code and Codex.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
GPT-5 mini 5.0%
Too close to call with Gemini 3.1 Flash Lite
tier 2, Harder tasks, tier 2 In Claude Code
Claude Haiku 4.5 4.2%
Too close to call with Claude Opus 5, Claude Sonnet 5, Claude Fable 5.1 and Claude Opus 5.5
tier 2, Harder tasks, tier 2
Some of these models come from another run of the same tasks In Codex
GPT-5.6 Terra 62.4%
No range published, so nobody is placed against it
tier 1, Codex run, tip by tip | Via API
This run does not publish a cost per task. In Claude Code
This run does not publish a cost per task. In Codex
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. | Via API
The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked. In Claude Code
Claude Opus 5, same answer on 8 of 10
Too close to call with Claude Haiku 4.5, Claude Fable 5.1 and Claude Sonnet 5
no sampling control here, Each tip asked five times In Codex
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
Gemini 3.1 Flash Lite 563 ms
Median time per call, across all twelve tips asked five times
2 models timed, Each tip asked five times In Claude Code
Claude Haiku 4.5 11.4 s
Median time per call, across all twelve tips asked five times
6 models timed, Each tip asked five times In Codex
No run of this job timed its calls this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | in another run | |
| GPT-5 mini | 5.0% tier 2 best | 8 of 10 at the vendor default |
| Gemini 3.1 Flash Lite | 3.7% tier 2 too close to call | 10 of 10 at temperature 0 |
| GPT-5.4 mini | in another run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 4.2% tier 2 best | 7 of 10 no sampling control here |
| Claude Fable 5 | in another run | |
| Claude Fable 5.1 | 0.3% tier 2 too close to call | 4 of 10 no sampling control here |
| Claude Opus 5 | 0.3% tier 2 too close to call | 8 of 10 no sampling control here |
| Claude Opus 5.5 | 0.3% tier 2 too close to call | 0 of 10 no sampling control here |
| Claude Sonnet 5 | 1.4% tier 2 too close to call | 7 of 10 no sampling control here |
In Codex
| Model | Score | Same answer twice |
|---|---|---|
| GPT-5.4 mini | 20.6% tier 1 | one pass, not measured |
| GPT-5.6 Luna | 57.7% tier 1 | one pass, not measured |
| GPT-5.6 Terra | 62.4% tier 1 best | one pass, not measured |
| GPT-6 Luna | 51.0% tier 1 | one pass, not measured |
| GPT-6 Sol | withheld: Incomplete coverage at retry exhaustion; 1 registered units have no valid record. Resume only within registered retry policy. | |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Getting the length right | GPT-5 mini | +86.5 percentage points (+76.7 to +96.3) | no one-line instruction was tested on tips |
| Getting the length right | Gemini 3.1 Flash Lite | +95.4 percentage points (+88.7 to +102.1) | no one-line instruction was tested on tips |
In Claude Code
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Getting the length right | Claude Opus 5 | +93.1 percentage points (+90.2 to +95.9) | no one-line instruction was tested on tips |
| Getting the length right | Claude Sonnet 5 | +93.3 percentage points (+88.2 to +98.5) | no one-line instruction was tested on tips |
| Getting the length right | Claude Haiku 4.5 | +77.4 percentage points (+67.9 to +86.9) | no one-line instruction was tested on tips |
| Getting the length right | Claude Fable 5.1 | +99.8 percentage points (+99.3 to +100.2) | no one-line instruction was tested on tips |
| Getting the length right | Claude Opus 5.5 | +99.8 percentage points (+99.3 to +100.2) | no one-line instruction was tested on tips |
In Codex
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Getting the length right | GPT-5.6 Terra | +66.0 percentage points (+45.1 to +86.9) | no one-line instruction was tested on tips |
| Getting the length right | GPT-5.6 Luna | +62.6 percentage points (+38.3 to +86.8) | no one-line instruction was tested on tips |
| Getting the length right | GPT-6 Sol | +71.7 percentage points (+51.2 to +92.1) | no one-line instruction was tested on tips |
| Getting the length right | GPT-6 Luna | +79.3 percentage points (+68.2 to +90.3) | no one-line instruction was tested on tips |
How hard the tasks were
The first tasks got too easy to tell the models apart, so we built a harder set, tier 2, and the scores here are from it. In Codex the harder set was not run, so those scores are from tier 1.
The full tables behind this page
Via API
Read from Harder tasks, tier 2, 2026-09-19.
- Another run, Main run, tip by tip, 2026-08-17: Gemini 3.1 Flash Lite 31.0% (tier 1), Claude Haiku 4.5 24.7% (tier 1), GPT-5 mini 1.6% (tier 1)
- Another run, API twin for the Codex run, 2026-09-07: GPT-5.4 mini 15.7% (tier 1)
In Claude Code
Read from Harder tasks, tier 2, 2026-09-19, with the same tasks also from Claude Opus 5.5, tier 2, 2026-09-25.
- Another run, Main run, 2026-09-03: Claude Fable 5 76.0% (tier 1), Claude Fable 5.1 50.2% (tier 1), Claude Haiku 4.5 27.5% (tier 1), Claude Sonnet 5 17.8% (tier 1), Claude Opus 5 15.0% (tier 1)
- Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 80.1% (tier 1)
- Another run, Claude Opus 5.5, asked each task five times, 2026-09-28: Claude Opus 5.5 76.6% (tier 1)
In Codex
Read from Codex run, tip by tip, 2026-10-03.
- Another run, Codex run, replaying two earlier tests, 2026-09-07: GPT-5.4 mini 18.6% (tier 1)
| Tip | Model | Test method | Run |
|---|---|---|---|
| Getting the length right | Claude Opus 5 | In Claude Code | tier 2, tier 2 |
| Getting the length right | Claude Sonnet 5 | In Claude Code | tier 2, tier 2 |
| Getting the length right | Claude Haiku 4.5 | In Claude Code | tier 2, tier 2 |
| Getting the length right | GPT-5 mini | Via API | tier 2, tier 2 |
| Getting the length right | Gemini 3.1 Flash Lite | Via API | tier 2, tier 2 |
| Getting the length right | Claude Fable 5.1 | In Claude Code | tier 2, tier 2 |
| Getting the length right | GPT-5.6 Terra | In Codex | tier 2, tier 2 |
| Getting the length right | GPT-5.6 Luna | In Codex | tier 2, tier 2 |
| Getting the length right | GPT-6 Sol | In Codex | tier 2, tier 2 |
| Getting the length right | GPT-6 Luna | In Codex | tier 2, tier 2 |
| Getting the length right | Claude Opus 5.5 | In Claude Code | Claude Opus 5.5, tier 2, tier 2 |
Version pairs on this job
- On getting the length right, Claude Fable 5.1 scores 50% unaided against 76% for Claude Fable 5; scores lower.
The job in full: Getting the length right. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/length-control.json.