Which LLM Admits Its Knowledge Cutoff? Tested
Answering about recent events means a question about something newer than the model knows, where the right answer says so. Each model was given the same tasks with no help, tested in the API, Claude Code and Codex.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
Claude Haiku 4.5 85.6%
tier 1, Main run, tip by tip In Claude Code
Claude Haiku 4.5 96.7%
Too close to call with Claude Sonnet 5, Claude Fable 5.1, Claude Fable 5 and Claude Opus 5
tier 1, Main run In Codex
Only one model was measured on this run, so there is nothing to rank. | Via API
The best model on this table publishes no cost per task, so a cheaper model cannot be priced against it. In Claude Code
Claude Haiku 4.5 $0.00220 a task
The best score is also the cheapest in range
run on subscription; shown at API list price for comparison In Codex
Only one model was measured on this run, so there is nothing to rank. | Via API
The models were asked under different settings, and a setting changes how often an answer repeats, so they are not ranked. In Claude Code
Claude Haiku 4.5, same answer on 28 of 30
Too close to call with Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5 and Claude Sonnet 5
no sampling control here, Each tip asked five times In Codex
No run asked these tasks more than once this way, so there is no repeat to compare. | Via API
Gemini 3.1 Flash Lite 563 ms
Median time per call, across all twelve tips asked five times
2 models timed, Each tip asked five times In Claude Code
Claude Haiku 4.5 11.4 s
Median time per call, across all twelve tips asked five times
6 models timed, Each tip asked five times In Codex
No run of this job timed its calls this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 85.6% tier 1 best | 28 of 30 temperature 0 asked for |
| GPT-5 mini | 48.6% tier 1 | 6 of 30 at the vendor default |
| Gemini 3.1 Flash Lite | 3.3% tier 1 | 30 of 30 at temperature 0 |
| GPT-5.4 mini | 27.5% tier 1 | 25 of 30 at temperature 0 |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 96.7% tier 1 best | 28 of 30 no sampling control here |
| Claude Fable 5 | 87.8% tier 1 too close to call | 20 of 30 no sampling control here |
| Claude Fable 5.1 | 92.2% tier 1 too close to call | 20 of 30 no sampling control here |
| Claude Opus 5 | 80.3% tier 1 too close to call | 23 of 30 no sampling control here |
| Claude Opus 5.5 | in another run | |
| Claude Sonnet 5 | 93.3% tier 1 too close to call | 22 of 30 no sampling control here |
In Codex
| Model | Score | Same answer twice |
|---|---|---|
| GPT-5.4 mini | withheld: model no longer served to this account after 38 of 60 units; two retries refused | |
| GPT-5.6 Luna | withheld: 6 units exhausted the single C8 retry after policy-blocked built-in tool attempts; no completed tool or MCP hit. Full one-pass coverage not reached. | |
| GPT-5.6 Terra | 30.0% tier 1 | one pass, not measured |
| GPT-6 Luna | withheld: Incomplete coverage at retry exhaustion; 9 registered units have no valid record. Resume only within registered retry policy. | |
| GPT-6 Sol | withheld: Incomplete coverage at retry exhaustion; 15 registered units have no valid record. Resume only within registered retry policy. | |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Questions about recent events | GPT-5 mini | +41.1 percentage points (+24.6 to +57.6) | no one-line instruction was tested on tips |
| Questions about recent events | Gemini 3.1 Flash Lite | +30.0 percentage points (+13.3 to +46.7) | no one-line instruction was tested on tips |
| Questions about recent events | GPT-5.4 mini | +59.2 percentage points (+41.0 to +77.4) | no one-line instruction was tested on tips |
In Codex
| Tip | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| Questions about recent events | GPT-5.4 mini | +54.2 percentage points (+37.9 to +70.4) | no one-line instruction was tested on tips |
| Model | Where | With no help |
|---|---|---|
| Claude Haiku 4.5 | In Claude Code | 96.7%, tier 1 |
| Claude Fable 5.1 | In Claude Code | 92.2%, tier 1 |
| Claude Sonnet 5 | In Claude Code | 93.3%, tier 1 |
Why the two test methods disagree here
| With no help | What the tip changed | |
|---|---|---|
| Via API | 85.6% | -23.3 points (-41.6 to -5.0) |
| In Claude Code | 96.7% | 2.5 points (-4.3 to 9.3) |
The tip hurt through the API and made no clear change in Claude Code. One likely reason: Claude Code tells the model the date before it answers. That is our guess, not something we have shown.
How hard the tasks were
These scores are from the first tasks we built for this job, tier 1, the set the models are read on.
The full tables behind this page
Via API
Read from Main run, tip by tip, 2026-08-17, with the same tasks also from API twin for the Codex run, 2026-09-07.
In Claude Code
Read from Main run, 2026-09-03.
- Another run, Opus 5.5 series, each task asked once, 2026-09-25: Claude Opus 5.5 100.0% (tier 1)
- Another run, Claude Opus 5.5, asked each task five times, 2026-09-28: Claude Opus 5.5 93.3% (tier 1)
In Codex
Read from Codex run, tip by tip, 2026-10-03.
- Another run, Codex run, replaying two earlier tests, 2026-09-07: GPT-5.4 mini 15.0% (tier 1)
| Tip | Model | Test method | Run |
|---|---|---|---|
| Questions about recent events | GPT-5 mini | Via API | Main run, tier 1 |
| Questions about recent events | Gemini 3.1 Flash Lite | Via API | Main run, tier 1 |
| Questions about recent events | GPT-5.4 mini | Via API | API twin for the Codex run, tier 1 |
| Questions about recent events | GPT-5.4 mini | In Codex | Codex run, replaying two earlier tests, tier 1 |
Version pairs on this job
- On questions about recent events, Claude Fable 5.1 scores 92% unaided against 88% for Claude Fable 5; no measurable change.
The job in full: Questions about recent events. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/recent-events.json.