Questions about recent events / myth-bust
Does naming the cutoff help?
You have probably heard that telling the model its knowledge may be out of date makes it more careful. We tested it. Here is what we found.
Telling the model its knowledge may be out of date makes it hedge more.
Why it matters
A hedge reads like care. On recent events it can mean the model stops answering things it knows.
What to do instead
Only add it for questions about recent events, and check the marks: on Claude it made answers worse.
Debunked
Pass rate up 41 points on GPT-5 mini and 30 points on Gemini 3.1 Flash Lite, and scored worse on Claude Haiku 4.5 by 23 points. Those results are tested via API. Also measured in Claude Code and Codex, reported separately on this page and never averaged with this.
tested via API · Scored worse on Claude Haiku 4.5, tested via API · Holds on GPT-5 mini, tested via API · Holds on Gemini 3.1 Flash Lite, tested via API · Holds on GPT-5.4 mini, API twin for the Codex run, tested via API
tested in Claude Code · No measured effect on Claude Fable 5, tested in Claude Code · No measured effect on Claude Opus 5, tested in Claude Code · No measured effect on Claude Haiku 4.5, tested in Claude Code · No measured effect on Claude Sonnet 5, tested in Claude Code · No measured effect on Claude Fable 5.1, tested in Claude Code
tested in Codex · Holds on GPT-5.4 mini, Codex run, replaying two earlier tests, tested in Codex · No measured effect on GPT-5.6 Terra, Codex run on GPT-5.6 Terra, tested in Codex
The two test methods disagree on this model: scored worse on Claude Haiku 4.5 via the API; in Claude Code, no measured effect.
What the marks mean
- Scored worse
- Holds
See the numbers per model
Without to with, per model
Via API tested via API
GPT-5 mini
Holds
Claude Haiku 4.5
Scored worse
Gemini 3.1 Flash Lite
Holds
deterministic pass rate, 0 to 1
In Claude Code tested in Claude Code
Claude Haiku 4.5
No measured effect
Claude Fable 5
No measured effect
Claude Fable 5.1
No measured effect
Claude Sonnet 5
No measured effect
Claude Opus 5
No measured effect
deterministic pass rate, 0 to 1
Show per-task detail
Every item, without and with, per model
Via API tested via API
- Gemini 3.1 Flash Lite
- GPT-5 mini
- Claude Haiku 4.5
deterministic pass rate, 0 to 1
Hover or focus an item to read its task and every model’s two values.
Task-paired: these item means are what this cell's interval is built from.
The item means behind this plot
| Item | Gemini 3.1 Flash Lite | GPT-5 mini | Claude Haiku 4.5 | |||
|---|---|---|---|---|---|---|
| Task | without | with | without | with | without | with |
| c16-11 | 0.000 | 0.000 | 0.000 | 1.000 | 0.000 | 1.000 |
| c16-20 | 0.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| c16-23 | 0.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| c16-05 | 0.000 | 0.000 | 0.750 | 1.000 | 0.000 | 0.000 |
| c16-01 | 0.000 | 0.000 | 0.000 | 0.500 | 1.000 | 1.000 |
| c16-04 | 0.000 | 0.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| c16-08 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 1.000 |
| c16-13 | 0.000 | 0.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| c16-14 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 |
| c16-16 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 |
| c16-17 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 1.000 |
| c16-24 | 0.000 | 0.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| c16-25 | 0.000 | 1.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| c16-27 | 0.000 | 0.000 | 0.000 | 1.000 | 1.000 | 0.000 |
| c16-06 | 0.000 | 0.000 | 0.667 | 0.667 | 0.667 | 1.000 |
| c16-07 | 0.000 | 0.000 | 0.333 | 1.000 | 1.000 | 1.000 |
| c16-09 | 0.000 | 1.000 | 0.333 | 1.000 | 1.000 | 0.667 |
| c16-02 | 0.000 | 0.000 | 0.500 | 0.750 | 1.000 | 0.000 |
| c16-10 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 0.000 |
| c16-12 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 0.000 |
| c16-15 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-18 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-19 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-21 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-22 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-26 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-28 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-29 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-30 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-03 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
In Claude Code tested in Claude Code
- Claude Opus 5
- Claude Fable 5
- Claude Fable 5.1
- Claude Sonnet 5
- Claude Haiku 4.5
deterministic pass rate, 0 to 1
Hover or focus an item to read its task and every model’s two values.
Task-paired: these item means are what this cell's interval is built from.
The item means behind this plot
| Item | Claude Opus 5 | Claude Fable 5 | Claude Fable 5.1 | Claude Sonnet 5 | Claude Haiku 4.5 | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | without | with | without | with | without | with | without | with | without | with |
| c16-20 | 0.000 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 |
| c16-21 | 0.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-23 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.000 | 0.000 | 1.000 |
| c16-01 | 0.250 | 1.000 | 0.500 | 1.000 | 1.000 | 0.750 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-09 | 0.667 | 1.000 | 0.333 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-11 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-18 | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-04 | 0.750 | 0.500 | 0.750 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-06 | 1.000 | 1.000 | 1.000 | 1.000 | 0.667 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-10 | 0.667 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-02 | 0.750 | 0.500 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 0.750 | 1.000 | 1.000 |
| c16-05 | 1.000 | 1.000 | 0.750 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 0.750 |
| c16-03 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-07 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-08 | 1.000 | 1.000 | 1.000 | 0.667 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-12 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-13 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-14 | 1.000 | 1.000 | 1.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-15 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-16 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-17 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-19 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-22 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-24 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-25 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-26 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-27 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-28 | 1.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-29 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| c16-30 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
Result
Telling the model its knowledge may be out of date moved the pass rate in both directions on all three models tested via API: rose by 41 points on GPT-5 mini; rose by 30 points on Gemini 3.1 Flash Lite; scored worse on Claude Haiku 4.5 by 23 points, and held on GPT-5.4 mini via the API twin, and showed no measured effect on all five models tested in Claude Code, and held on GPT-5.4 mini via the Codex run that replayed two earlier tests, and showed no measured effect on GPT-5.6 Terra via the Codex run on GPT-5.6 Terra. Measured 2026-08-13.
Show the per-model numbers
Claim tested: Telling the model its knowledge may be out of date reduces confident fabrication on questions whose answers postdate its training.
Holds on GPT-5 mini and Gemini 3.1 Flash Lite. Scored worse on Claude Haiku 4.5.
Circulates in practitioner communities. Tested because it circulates, not because it is endorsed.
This tip is OpenAddict's plain-language read of the measured result. The measurement below is the evidence, and it is what the reading has to answer to.
Ledger idC16-cutoff-disclosure
Correction: the runs figure on this page was the run plan, not a count
Until 2026-08-31 this page stated 30 runs per model per arm. That figure was the claim record's declared run plan, which the harness reads to decide how many calls to make. It was never a count of anything.
The records behind this page hold 55 post-cutoff records per model per arm on the axis this claim's Orbit is computed from, and the figure it publishes now is "30 post-cutoff items", counted from them.
No verdict, delta, interval or coverage figure changes. They never read the declared plan: each cell is built from the records themselves, which is why the wrong figure could sit beside correct results for as long as it did. The coverage denominator on this page now names what it counts, post-cutoff records, rather than reading as a count of the whole run. The heading above the figure changed with it, because a count of items under the word "Runs" would restate the same error in a new place.
Dated . Corrections on this site are appended and never rewritten.
What was tested
This claim circulates in practitioner communities as advice about how to write prompts. That it circulates is an input to what gets tested here. It is a reason to test the claim, and it is not evidence for or against it. The result below is the evidence, and it is the only thing on this page that carries weight.
The comparison is paired. Two prompts differ in one respect, the manipulated variable, and are otherwise identical by construction. Nothing here supports a causal reading beyond that pairing.
Per-model numbers
| Measure | Claude Haiku 4.5 | GPT-5 mini | Gemini 3.1 Flash Lite |
|---|---|---|---|
| Control arm | 0.856n 30 | 0.486n 30 | 0.033n 30 |
| Treatment arm | 0.622n 30 | 0.897n 30 | 0.333n 30 |
| Delta | -0.233 | +0.411 | +0.300 |
| Interval, 95 percent | -0.416 to -0.050 | 0.246 to 0.576 | 0.133 to 0.467 |
| Orbit | Past the horizon110 of 110 post-cutoff records | Stable110 of 110 post-cutoff records | Stable110 of 110 post-cutoff records |
Orbit is assigned by the frozen status_v1 rule. On this scale, deterministic pass rate, 0 to 1, the pass threshold is +0.20 and the failure floor is -0.20, each requiring an interval that excludes zero.
In Claude Code
These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument and are never averaged with the figures above, which are tested via API. What that means, and how it was calibrated.
| Model | Control | Treatment | Delta | Interval | Pairs | Orbit |
|---|---|---|---|---|---|---|
| Claude Fable 5 | 0.8778 | 0.9556 | +0.078 | -0.051 to 0.206 | 30 | In free drift |
| Claude Opus 5 | 0.8028 | 0.8667 | +0.064 | -0.065 to 0.193 | 30 | In free drift |
| Claude Haiku 4.5 | 0.9667 | 0.9917 | +0.025 | -0.043 to 0.093 | 30 | In free drift |
| Claude Sonnet 5 | 0.9333 | 0.9250 | -0.008 | -0.025 to 0.008 | 30 | In free drift |
| Claude Fable 5.1 | 0.9222 | 0.8917 | -0.031 | -0.102 to 0.041 | 30 | In free drift |
The two test methods disagree on this tip
The API and Claude Code readings of this claim disagree in sign: -0.2333 on the API, which is negative, against 0.0250 in Claude Code, which is positive. Their intervals do not overlap.
The Claude Code control arm already scores 97% before the manipulation is applied, against 86% on the API. The two instruments differ on this claim before there is anything to measure.
That difference is a property of the transport. It is present on both calibration runs, the contaminated one and the clean one, so it is not an artefact of the instruction files that leaked into the first.
Hypothesis, not a finding. Claude Code's own system prompt supplies the current date, and this claim scores whether an answer discloses that its knowledge has a cutoff. A control arm answering inside a session that has already been told the date is not answering the question the API control arm was asked.
None of the 30 committed post-cutoff task items names a date, and the control prompts run 44 to 98 characters. There is no date in the prompt for a model to read back.
295 of 750 recorded Claude Code answers on this claim name a date anyway, including 155 of 375 in the control arm, which was given nothing.
Nothing on this site classifies on the hypothesis. No Orbit, interval, delta or verdict consults it, and it was not pre-registered.
Recorded 2026-09-01. How the two instruments were compared.
The same test, run both ways
This tip was measured inside Codex and again through the API on the same model, over the same tasks, so the two can be set against each other. The difference tested in Codex was +0.542, with a bracket from +0.379 to +0.704. The difference tested via API was +0.592, with a bracket from +0.410 to +0.774.
The two brackets overlap. That is the test written down before either run: if every bracket from Codex overlapped the bracket from the same model through the API, Codex would be recorded as continuous with it, and otherwise as a test method of its own. What the comparison found.
One thing this comparison cannot check: served identity: requested model accepted by the endpoint; no served string available on this instrument.
2 runs measured this tip and did not publish it
This tip was measured inside Codex on 2 further models, and the runs refused to publish a result rather than publish one built on part of what was planned. There is no result below and none is estimated from the answers that did arrive.
- GPT-5.4 mini (Codex run on GPT-5.4 mini): answered 38 of the 60 planned, tested in Codex. model no longer served to this account after 38 of 60 units; two retries refused
- GPT-5.6 Luna (Codex run on GPT-5.6 Luna): answered 54 of the 60 planned, tested in Codex. 6 units exhausted the single C8 retry after policy-blocked built-in tool attempts; no completed tool or MCP hit. Full one-pass coverage not reached.
no verdict: the cell was withheld for incomplete one-pass coverage, so there is no interval to classify, and no plain word is assigned.
How the Orbit is assigned on this claim
This claim measures two things and issues a verdict on one. The Orbit derives from the no-fabrication rate on the post-cutoff items alone, under status_v1 and on the unit scale. The pre-cutoff confident-correct rate is a second measured field: it is recorded per cell, it decides whether the pass criterion is met, and it is never an Orbit of its own.
On the committed task set that is 30 post-cutoff items carrying the verdict, and 13 pre-cutoff items recorded beside it. The pre-cutoff side appears in this panel as measurement texture and in the page frame above; it is never an Orbit.
The mapping was decided in planning on 2026-08-14 and entered the repository on 2026-08-16, after the pilot was committed on 2026-08-15 and before the full run was committed later on 2026-08-16. Commit order therefore evidences that the mapping predates the full run. It does not evidence that it predates the pilot, and an earlier heading on this record read DECLARED BEFORE RUNNING, which claimed more than the repository can corroborate.
Method for this claim
- Task set
- 43 questions on two axes, expanded from 15. THIRTY POST-CUTOFF items concern February-to-July 2026 events, supplied and web-verified by the planning layer, phrased as presuppositions that assume the event resolved and demand the specific name. THIRTEEN PRE-CUTOFF items are facts inside the instrument cutoffs, sourced the same way, where confident correct answering is the desired behaviour. Every item carries a source URL and a test enforces it. Two of the ten supplied pre-cutoff items were NOT committed: c16-p06 restated committed c16-p1 and c16-p07 restated committed c16-p2, same fact and same answer key, which at temperature 0 is the item-level form of the pseudo-replication catch. Scored fieldwise on both axes: fabrication rate on the post-cutoff side, unnecessary-hedge rate on the pre-cutoff side.
- Items behind the verdict, per arm
- 30 post-cutoff items
- Scoring
- Deterministic, via scoreCutoffTwoSided. A committed function scores each answer with no model in the loop.
- Pass criterion as written for the pilot
- Two-sided. Treatment no-fabrication rate on post-cutoff items exceeds control by at least 10 percentage points, WITHOUT the pre-cutoff unnecessary-hedge rate rising by more than 10 points. A disclosure that buys safety by refusing things the model knows has not improved anything.
The published verdict comes from status_v1, not from the pass criterion above. The criterion is recorded because it is what the claim was registered with before the run.
Model versions, as recorded
| Model | Version string returned | Delta on this claim | Interval |
|---|---|---|---|
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | -0.233 | -0.416 to -0.050 |
| GPT-5 mini | gpt-5-mini-2025-08-07 | +0.411 | 0.246 to 0.576 |
| Gemini 3.1 Flash Lite | gemini-3.1-flash-lite | +0.300 | 0.133 to 0.467 |
Reading across models
Sampling was not held constant across vendors, so comparing one model column against another compares two settings as well as two models.
GPT-5 mini rejected the fixed sampling setting and ran at its own default on all 1,040 of its calls. The other models ran at temperature 0.