Which LLM Returns Clean JSON? Tested
Getting clean JSON back means asking for data in an exact format and getting nothing else in the reply. Each model was given the same ten tasks with no help, tested in the API, Claude Code and Codex.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. | Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. | Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. | Via API
No model managed this on its own. In Claude Code
No model managed this on its own. In Codex
No model managed this on its own. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | in another run | |
| GPT-5 mini | 0.0% tier 2 | 10 of 10 at the vendor default |
| Gemini 3.1 Flash Lite | 0.0% tier 2 | 10 of 10 at temperature 0 |
| GPT-5.4 mini | not run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 0.0% tier 2 | 10 of 10 no sampling control here |
| Claude Fable 5 | in another run | |
| Claude Fable 5.1 | 0.0% tier 2 | 10 of 10 no sampling control here |
| Claude Opus 5 | 0.0% tier 2 | 10 of 10 no sampling control here |
| Claude Opus 5.5 | 0.0% tier 2 | 10 of 10 no sampling control here |
| Claude Sonnet 5 | 0.0% tier 2 | 10 of 10 no sampling control here |
In Codex
| Model | Score | Same answer twice |
|---|---|---|
| GPT-5.4 mini | 0.0% tier 1 | one pass, not measured |
| GPT-5.6 Luna | 0.0% tier 1 | one pass, not measured |
| GPT-5.6 Terra | 0.0% tier 1 | one pass, not measured |
| GPT-6 Luna | 0.0% tier 1 | one pass, not measured |
| GPT-6 Sol | 0.0% tier 1 | one pass, not measured |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Getting clean JSON back | GPT-5 mini | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
| Getting clean JSON back | Gemini 3.1 Flash Lite | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
In Claude Code
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Getting clean JSON back | Claude Opus 5 | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
| Getting clean JSON back | Claude Sonnet 5 | +80.0 percentage points (+53.9 to +106.1) | no one-line instruction was tested on tips |
| Getting clean JSON back | Claude Haiku 4.5 | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
| Getting clean JSON back | Claude Fable 5.1 | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
| Getting clean JSON back | Claude Opus 5.5 | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
In Codex
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Getting clean JSON back | GPT-5.6 Terra | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
| Getting clean JSON back | GPT-5.6 Luna | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
| Getting clean JSON back | GPT-6 Sol | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
| Getting clean JSON back | GPT-6 Luna | +100.0 percentage points (+100.0 to +100.0) | no one-line instruction was tested on tips |
How hard the tasks were
The first tasks got too easy to tell the models apart, so we built a harder set, tier 2, and the scores here are from it. In Codex the harder set was not run, so those scores are from tier 1.
The full tables behind this page
Via API
Read from Harder tasks, tier 2, 2026-09-19.
- Another run, Main run, tip by tip, 2026-08-17: Claude Haiku 4.5 0.0% (tier 1), Gemini 3.1 Flash Lite 0.0% (tier 1), GPT-5 mini 0.0% (tier 1)
In Claude Code
Read from Harder tasks, tier 2, 2026-09-19, with the same tasks also from Claude Opus 5.5, tier 2, 2026-09-25.
- Another run, Main run, 2026-09-03: Claude Fable 5 0.0% (tier 1), Claude Fable 5.1 0.0% (tier 1), Claude Haiku 4.5 0.0% (tier 1), Claude Opus 5 0.0% (tier 1), Claude Sonnet 5 0.0% (tier 1)
- Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 0.0% (tier 1)
- Another run, Claude Opus 5.5, asked each task five times, 2026-09-28: Claude Opus 5.5 0.0% (tier 1)
In Codex
Read from Codex run, tip by tip, 2026-10-03.
| Tip | Model | Test method | Run |
|---|---|---|---|
| Getting clean JSON back | Claude Opus 5 | In Claude Code | tier 2, tier 2 |
| Getting clean JSON back | Claude Sonnet 5 | In Claude Code | tier 2, tier 2 |
| Getting clean JSON back | Claude Haiku 4.5 | In Claude Code | tier 2, tier 2 |
| Getting clean JSON back | GPT-5 mini | Via API | tier 2, tier 2 |
| Getting clean JSON back | Gemini 3.1 Flash Lite | Via API | tier 2, tier 2 |
| Getting clean JSON back | Claude Fable 5.1 | In Claude Code | tier 2, tier 2 |
| Getting clean JSON back | GPT-5.6 Terra | In Codex | tier 2, tier 2 |
| Getting clean JSON back | GPT-5.6 Luna | In Codex | tier 2, tier 2 |
| Getting clean JSON back | GPT-6 Sol | In Codex | tier 2, tier 2 |
| Getting clean JSON back | GPT-6 Luna | In Codex | tier 2, tier 2 |
| Getting clean JSON back | Claude Opus 5.5 | In Claude Code | Claude Opus 5.5, tier 2, tier 2 |
Version pairs on this job
- On getting clean JSON back, Claude Fable 5.1 scores 0% unaided against 0% for Claude Fable 5; could not be measured.
The job in full: Getting clean JSON back. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/json-extraction.json.