Best LLM for Data Checking, Tested

Data checking means an agent checking a set of published statements against their sources with tools, and fixing the wrong ones. Each model ran it five times in Claude Code.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
In Claude Code Claude Opus 5 75.8% No range published, so nobody is placed against it tier 1, Workflow run, five times each
In Claude Code This run does not publish a cost per task.
In Claude Code No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code No run of this job timed its calls this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.5not run: Before we buy a single run we check that a session can only read the files we put in front of it. That check never came back the way it has to, on any of the 3 tries we had written down beforehand, so we bought no runs at all. This is us declining to measure, and it says nothing about the model.
Claude Fable 5not run
Claude Fable 5.171.7% tier 1not asked twice
Claude Opus 575.8% tier 1 bestnot asked twice
Claude Opus 5.5not run
Claude Sonnet 564.2% tier 1not asked twice

What measurably helps on this job

No tip or skill has been tested on this job yet.

When an input was missing

One input was left out on purpose. In one run nothing said so. In the other, the model was told.

ModelNot toldTold
Claude Opus 5stopped 5 of 5, went ahead 0 of 5stopped 5 of 5, went ahead 0 of 5
Claude Sonnet 5stopped 5 of 5, went ahead 0 of 5stopped 5 of 5, went ahead 0 of 5
Claude Haiku 4.5stopped 2 of 4, went ahead 2 of 4not run
Claude Fable 5.1stopped 4 of 4, went ahead 0 of 4stopped 5 of 5, went ahead 0 of 5
Claude Opus 5.5stopped 5 of 5, went ahead 0 of 5not run

How hard the tasks were

A workflow has no tiers. Each model ran the same steps on the same material, and one input was left out on purpose.

The full tables behind this page

In Claude Code

Read from Workflow run, five times each, 2026-09-10.

The job in full: Corpus integrity and correction. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/corpus-integrity.json.