Which LLM Returns Clean JSON? Tested

Getting clean JSON back means asking for data in an exact format and getting nothing else in the reply. Each model was given the same ten tasks with no help, tested in the API, Claude Code and Codex.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.
Via API No model managed this on its own.
In Claude Code No model managed this on its own.
In Codex No model managed this on its own.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.5in another run
GPT-5 mini0.0% tier 210 of 10 at the vendor default
Gemini 3.1 Flash Lite0.0% tier 210 of 10 at temperature 0
GPT-5.4 mininot run

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.50.0% tier 210 of 10 no sampling control here
Claude Fable 5in another run
Claude Fable 5.10.0% tier 210 of 10 no sampling control here
Claude Opus 50.0% tier 210 of 10 no sampling control here
Claude Opus 5.50.0% tier 210 of 10 no sampling control here
Claude Sonnet 50.0% tier 210 of 10 no sampling control here

In Codex

ModelScoreSame answer twice
GPT-5.4 mini0.0% tier 1one pass, not measured
GPT-5.6 Luna0.0% tier 1one pass, not measured
GPT-5.6 Terra0.0% tier 1one pass, not measured
GPT-6 Luna0.0% tier 1one pass, not measured
GPT-6 Sol0.0% tier 1one pass, not measured

What measurably helps on this job

Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.

Via API

TipModelBetter than no help byAgainst one plain sentence
Getting clean JSON backGPT-5 mini+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips
Getting clean JSON backGemini 3.1 Flash Lite+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips

In Claude Code

TipModelBetter than no help byAgainst one plain sentence
Getting clean JSON backClaude Opus 5+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips
Getting clean JSON backClaude Sonnet 5+80.0 percentage points (+53.9 to +106.1)no one-line instruction was tested on tips
Getting clean JSON backClaude Haiku 4.5+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips
Getting clean JSON backClaude Fable 5.1+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips
Getting clean JSON backClaude Opus 5.5+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips

In Codex

TipModelBetter than no help byAgainst one plain sentence
Getting clean JSON backGPT-5.6 Terra+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips
Getting clean JSON backGPT-5.6 Luna+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips
Getting clean JSON backGPT-6 Sol+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips
Getting clean JSON backGPT-6 Luna+100.0 percentage points (+100.0 to +100.0)no one-line instruction was tested on tips

How hard the tasks were

The first tasks got too easy to tell the models apart, so we built a harder set, tier 2, and the scores here are from it. In Codex the harder set was not run, so those scores are from tier 1.

The full tables behind this page

Via API

Read from Harder tasks, tier 2, 2026-09-19.

  • Another run, Main run, tip by tip, 2026-08-17: Claude Haiku 4.5 0.0% (tier 1), Gemini 3.1 Flash Lite 0.0% (tier 1), GPT-5 mini 0.0% (tier 1)

In Claude Code

Read from Harder tasks, tier 2, 2026-09-19, with the same tasks also from Claude Opus 5.5, tier 2, 2026-09-25.

  • Another run, Main run, 2026-09-03: Claude Fable 5 0.0% (tier 1), Claude Fable 5.1 0.0% (tier 1), Claude Haiku 4.5 0.0% (tier 1), Claude Opus 5 0.0% (tier 1), Claude Sonnet 5 0.0% (tier 1)
  • Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 0.0% (tier 1)
  • Another run, Claude Opus 5.5, asked each task five times, 2026-09-28: Claude Opus 5.5 0.0% (tier 1)

In Codex

Read from Codex run, tip by tip, 2026-10-03.

Every result above, with the run it was read from.
TipModelTest methodRun
Getting clean JSON backClaude Opus 5In Claude Codetier 2, tier 2
Getting clean JSON backClaude Sonnet 5In Claude Codetier 2, tier 2
Getting clean JSON backClaude Haiku 4.5In Claude Codetier 2, tier 2
Getting clean JSON backGPT-5 miniVia APItier 2, tier 2
Getting clean JSON backGemini 3.1 Flash LiteVia APItier 2, tier 2
Getting clean JSON backClaude Fable 5.1In Claude Codetier 2, tier 2
Getting clean JSON backGPT-5.6 TerraIn Codextier 2, tier 2
Getting clean JSON backGPT-5.6 LunaIn Codextier 2, tier 2
Getting clean JSON backGPT-6 SolIn Codextier 2, tier 2
Getting clean JSON backGPT-6 LunaIn Codextier 2, tier 2
Getting clean JSON backClaude Opus 5.5In Claude CodeClaude Opus 5.5, tier 2, tier 2

Version pairs on this job

  • On getting clean JSON back, Claude Fable 5.1 scores 0% unaided against 0% for Claude Fable 5; could not be measured.

The job in full: Getting clean JSON back. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/json-extraction.json.