Best LLM for Data Checking, Tested
Data checking means an agent checking a set of published statements against their sources with tools, and fixing the wrong ones. Each model ran it five times in Claude Code.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
In Claude Code
Claude Opus 5 75.8%
No range published, so nobody is placed against it
tier 1, Workflow run, five times each | In Claude Code
This run does not publish a cost per task. | In Claude Code
No run asked these tasks more than once this way, so there is no repeat to compare. | In Claude Code
No run of this job timed its calls this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | not run: Before we buy a single run we check that a session can only read the files we put in front of it. That check never came back the way it has to, on any of the 3 tries we had written down beforehand, so we bought no runs at all. This is us declining to measure, and it says nothing about the model. | |
| Claude Fable 5 | not run | |
| Claude Fable 5.1 | 71.7% tier 1 | not asked twice |
| Claude Opus 5 | 75.8% tier 1 best | not asked twice |
| Claude Opus 5.5 | not run | |
| Claude Sonnet 5 | 64.2% tier 1 | not asked twice |
What measurably helps on this job
No tip or skill has been tested on this job yet.
When an input was missing
One input was left out on purpose. In one run nothing said so. In the other, the model was told.
| Model | Not told | Told |
|---|---|---|
| Claude Opus 5 | stopped 5 of 5, went ahead 0 of 5 | stopped 5 of 5, went ahead 0 of 5 |
| Claude Sonnet 5 | stopped 5 of 5, went ahead 0 of 5 | stopped 5 of 5, went ahead 0 of 5 |
| Claude Haiku 4.5 | stopped 2 of 4, went ahead 2 of 4 | not run |
| Claude Fable 5.1 | stopped 4 of 4, went ahead 0 of 4 | stopped 5 of 5, went ahead 0 of 5 |
| Claude Opus 5.5 | stopped 5 of 5, went ahead 0 of 5 | not run |
How hard the tasks were
A workflow has no tiers. Each model ran the same steps on the same material, and one input was left out on purpose.
The full tables behind this page
In Claude Code
Read from Workflow run, five times each, 2026-09-10.
The job in full: Corpus integrity and correction. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/corpus-integrity.json.