Which LLM Admits Its Knowledge Cutoff? Tested

Answering about recent events means a question about something newer than the model knows, where the right answer says so. Each model was given the same tasks with no help, tested in the API, Claude Code and Codex.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API Claude Haiku 4.5 85.6% tier 1, Main run, tip by tip
In Claude Code Claude Haiku 4.5 96.7% Too close to call with Claude Sonnet 5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 tier 1, Main run
In Codex Only one model was measured on this run, so there is nothing to rank.
Via API The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it.
In Claude Code Claude Haiku 4.5 $0.00220 a task The best score is also the cheapest in range run on subscription; shown at API list price for comparison
In Codex Only one model was measured on this run, so there is nothing to rank.
Via API The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked.
In Claude Code Claude Haiku 4.5, same answer on 28 of 30 Too close to call with Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5 and Claude Sonnet 5 no sampling control here, Each tip asked five times
In Codex No run asked these tasks more than once this way, so there is no repeat to compare.
Via API Gemini 3.1 Flash Lite 563 ms Median time per call, across all twelve tips asked five times 2 models timed, Each tip asked five times
In Claude Code Claude Haiku 4.5 11.4 s Median time per call, across all twelve tips asked five times 6 models timed, Each tip asked five times
In Codex No run of this job timed its calls this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.585.6% tier 1 best28 of 30 temperature 0 asked for
GPT-5 mini48.6% tier 16 of 30 at the vendor default
Gemini 3.1 Flash Lite3.3% tier 130 of 30 at temperature 0
GPT-5.4 mini27.5% tier 125 of 30 at temperature 0

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.596.7% tier 1 best28 of 30 no sampling control here
Claude Fable 587.8% tier 1 too close to call20 of 30 no sampling control here
Claude Fable 5.192.2% tier 1 too close to call20 of 30 no sampling control here
Claude Opus 580.3% tier 1 too close to call23 of 30 no sampling control here
Claude Opus 5.5in another run
Claude Sonnet 593.3% tier 1 too close to call22 of 30 no sampling control here

In Codex

ModelScoreSame answer twice
GPT-5.4 miniwithheld: model no longer served to this account after 38 of 60 units; two retries refused
GPT-5.6 Lunawithheld: 6 units exhausted the single C8 retry after policy-blocked built-in tool attempts; no completed tool or MCP hit. Full one-pass coverage not reached.
GPT-5.6 Terra30.0% tier 1one pass, not measured
GPT-6 Lunawithheld: Incomplete coverage at retry exhaustion; 9 registered units have no valid record. Resume only within registered retry policy.
GPT-6 Solwithheld: Incomplete coverage at retry exhaustion; 15 registered units have no valid record. Resume only within registered retry policy.

What measurably helps on this job

Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.

Via API

TipModelBetter than no help byAgainst one plain sentence
Questions about recent eventsGPT-5 mini+41.1 percentage points (+24.6 to +57.6)no one-line instruction was tested on tips
Questions about recent eventsGemini 3.1 Flash Lite+30.0 percentage points (+13.3 to +46.7)no one-line instruction was tested on tips
Questions about recent eventsGPT-5.4 mini+59.2 percentage points (+41.0 to +77.4)no one-line instruction was tested on tips

In Codex

TipModelBetter than no help byAgainst one plain sentence
Questions about recent eventsGPT-5.4 mini+54.2 percentage points (+37.9 to +70.4)no one-line instruction was tested on tips
Nothing left to measure: these models already scored near the top with no help.
ModelWhereWith no help
Claude Haiku 4.5In Claude Code96.7%, tier 1
Claude Fable 5.1In Claude Code92.2%, tier 1
Claude Sonnet 5In Claude Code93.3%, tier 1

Why the two test methods disagree here

 With no helpWhat the tip changed
Via API85.6%-23.3 points (-41.6 to -5.0)
In Claude Code96.7%2.5 points (-4.3 to 9.3)

The tip hurt through the API and made no clear change in Claude Code. One likely reason: Claude Code tells the model the date before it answers. That is our guess, not something we have shown.

How hard the tasks were

These scores are from the first tasks we built for this job, tier 1, the set the models are read on.

The full tables behind this page

Via API

Read from Main run, tip by tip, 2026-08-17, with the same tasks also from API twin for the Codex run, 2026-09-07.

In Claude Code

Read from Main run, 2026-09-03.

  • Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 100.0% (tier 1)
  • Another run, Claude Opus 5.5, asked each task five times, 2026-09-28: Claude Opus 5.5 93.3% (tier 1)

In Codex

Read from Codex run, tip by tip, 2026-10-03.

  • Another run, Codex run, replaying two earlier tests, 2026-09-07: GPT-5.4 mini 15.0% (tier 1)
Every result above, with the run it was read from.
TipModelTest methodRun
Questions about recent eventsGPT-5 miniVia APIMain run, tier 1
Questions about recent eventsGemini 3.1 Flash LiteVia APIMain run, tier 1
Questions about recent eventsGPT-5.4 miniVia APIAPI twin for the Codex run, tier 1
Questions about recent eventsGPT-5.4 miniIn CodexCodex run, replaying two earlier tests, tier 1

Version pairs on this job

  • On questions about recent events, Claude Fable 5.1 scores 92% unaided against 88% for Claude Fable 5; no measurable change.

The job in full: Questions about recent events. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/recent-events.json.