Best LLM for Writing Specs, Tested
Writing a spec means turning a product idea into a written plan with the parts and the tests it is asked for. Each model was given the same ten tasks with no help, tested in the API and Claude Code.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
Claude Haiku 4.5 99.5%
Too close to call with Gemini 3.1 Flash Lite
tier 1, First run, with and without the skill In Claude Code
Claude Opus 5.5 99.5%
Too close to call with Claude Haiku 4.5 and Claude Fable 5.1
tier 1, Wider run, three ways
Some of these models come from another run of the same tasks | Via API
Gemini 3.1 Flash Lite $0.00080 a task
Against Claude Haiku 4.5 $0.00332 a task, the best score
billed through the API In Claude Code
Claude Haiku 4.5 $0.00315 a task
Against Claude Opus 5.5 $0.04784 a task, the best score
run on subscription; shown at API list price for comparison
Claude Opus 5.5 is from Opus 5.5 series, wider run, three ways | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 27 of 40
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 99.5% tier 1 best | one pass, not measured |
| GPT-5 mini | 94.5% tier 1 | one pass, not measured |
| Gemini 3.1 Flash Lite | 99.1% tier 1 too close to call | one pass, not measured |
| GPT-5.4 mini | not run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 99.2% tier 1 too close to call | one pass, not measured |
| Claude Fable 5 | 95.9% tier 1 | one pass, not measured |
| Claude Fable 5.1 | 97.9% tier 1 too close to call | one pass, not measured |
| Claude Opus 5 | 92.3% tier 1 | one pass, not measured |
| Claude Opus 5.5 | 99.5% tier 1 best | one pass, not measured |
| Claude Sonnet 5 | 96.2% tier 1 | one pass, not measured |
What measurably helps on this job
Nothing we tested here helped by more than chance.
| Model | Where | With no help |
|---|---|---|
| Claude Haiku 4.5 | Via API | 99.5%, tier 1 |
| GPT-5 mini | Via API | 94.5%, tier 1 |
| Gemini 3.1 Flash Lite | Via API | 99.1%, tier 1 |
| Claude Haiku 4.5 | In Claude Code | 99.2%, tier 1 |
| Claude Fable 5 | In Claude Code | 95.9%, tier 1 |
| Claude Fable 5.1 | In Claude Code | 97.9%, tier 1 |
| Claude Opus 5 | In Claude Code | 92.3%, tier 1 |
| Claude Opus 5.5 | In Claude Code | 99.5%, tier 1 |
| Claude Sonnet 5 | In Claude Code | 96.2%, tier 1 |
How hard the tasks were
These scores are from the first tasks we built for this job, tier 1, the set the models are read on.
The full tables behind this page
Via API
Read from First run, with and without the skill, 2026-08-30.
- Another run, Wider run, three ways, 2026-09-06: Gemini 3.1 Flash Lite 99.1% (tier 1), GPT-5 mini 96.4% (tier 1)
In Claude Code
Read from Wider run, three ways, 2026-09-06, with the same tasks also from Opus 5.5 series, wider run, three ways, 2026-09-26; Fable 5, wider run, three ways, 2026-09-19.
- Another run, Second run, every task twice, 2026-08-31: Claude Haiku 4.5 98.6% (tier 1), Claude Fable 5 95.7% (tier 1), Claude Sonnet 5 94.8% (tier 1), Claude Opus 5 92.5% (tier 1)
- Another run, First run, single attempt, 2026-08-31: Claude Fable 5 93.2% (tier 1), Claude Opus 5 92.3% (tier 1)
- Another run, Fable 5.1, single attempt, 2026-09-03: Claude Fable 5.1 97.7% (tier 1)
Version pairs on this job
- On spec writing, Claude Fable 5.1 finds 98% of required sections and acceptance criteria unaided against 93% for Claude Fable 5; scores higher.
- On spec writing, wider run, Claude Fable 5.1 finds 98% of required sections and acceptance criteria unaided against 96% for Claude Fable 5; scores higher.
The job in full: Spec writing. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/spec-writing.json.