Best LLM for Writing Specs, Tested

Writing a spec means turning a product idea into a written plan with the parts and the tests it is asked for. Each model was given the same ten tasks with no help, tested in the API and Claude Code.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API Claude Haiku 4.5 99.5% Too close to call with Gemini 3.1 Flash Lite tier 1, First run, with and without the skill
In Claude Code Claude Opus 5.5 99.5% Too close to call with Claude Haiku 4.5 and Claude Fable 5.1 tier 1, Wider run, three ways Some of these models come from another run of the same tasks
Via API Gemini 3.1 Flash Lite $0.00080 a task Against Claude Haiku 4.5 $0.00332 a task, the best score billed through the API
In Claude Code Claude Haiku 4.5 $0.00315 a task Against Claude Opus 5.5 $0.04784 a task, the best score run on subscription; shown at API list price for comparison Claude Opus 5.5 is from Opus 5.5 series, wider run, three ways
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code Claude Opus 5, same answer on 27 of 40 no sampling control here, Second run, every task twice
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.599.5% tier 1 bestone pass, not measured
GPT-5 mini94.5% tier 1one pass, not measured
Gemini 3.1 Flash Lite99.1% tier 1 too close to callone pass, not measured
GPT-5.4 mininot run

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.599.2% tier 1 too close to callone pass, not measured
Claude Fable 595.9% tier 1one pass, not measured
Claude Fable 5.197.9% tier 1 too close to callone pass, not measured
Claude Opus 592.3% tier 1one pass, not measured
Claude Opus 5.599.5% tier 1 bestone pass, not measured
Claude Sonnet 596.2% tier 1one pass, not measured

What measurably helps on this job

Nothing we tested here helped by more than chance.

Nothing left to measure: these models already scored near the top with no help.
ModelWhereWith no help
Claude Haiku 4.5Via API99.5%, tier 1
GPT-5 miniVia API94.5%, tier 1
Gemini 3.1 Flash LiteVia API99.1%, tier 1
Claude Haiku 4.5In Claude Code99.2%, tier 1
Claude Fable 5In Claude Code95.9%, tier 1
Claude Fable 5.1In Claude Code97.9%, tier 1
Claude Opus 5In Claude Code92.3%, tier 1
Claude Opus 5.5In Claude Code99.5%, tier 1
Claude Sonnet 5In Claude Code96.2%, tier 1

How hard the tasks were

These scores are from the first tasks we built for this job, tier 1, the set the models are read on.

The full tables behind this page

Via API

Read from First run, with and without the skill, 2026-08-30.

  • Another run, Wider run, three ways, 2026-09-06: Gemini 3.1 Flash Lite 99.1% (tier 1), GPT-5 mini 96.4% (tier 1)

In Claude Code

Read from Wider run, three ways, 2026-09-06, with the same tasks also from Opus 5.5 series, wider run, three ways, 2026-09-26; Fable 5, wider run, three ways, 2026-09-19.

  • Another run, Second run, every task twice, 2026-08-31: Claude Haiku 4.5 98.6% (tier 1), Claude Fable 5 95.7% (tier 1), Claude Sonnet 5 94.8% (tier 1), Claude Opus 5 92.5% (tier 1)
  • Another run, First run, single attempt, 2026-08-31: Claude Fable 5 93.2% (tier 1), Claude Opus 5 92.3% (tier 1)
  • Another run, Fable 5.1, single attempt, 2026-09-03: Claude Fable 5.1 97.7% (tier 1)

Version pairs on this job

  • On spec writing, Claude Fable 5.1 finds 98% of required sections and acceptance criteria unaided against 93% for Claude Fable 5; scores higher.
  • On spec writing, wider run, Claude Fable 5.1 finds 98% of required sections and acceptance criteria unaided against 96% for Claude Fable 5; scores higher.

The job in full: Spec writing. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/spec-writing.json.