Questions about recent events / myth-bust

Does naming the cutoff help?

You have probably heard that telling the model its knowledge may be out of date makes it more careful. We tested it. Here is what we found.

Telling the model its knowledge may be out of date makes it hedge more.

Why it matters

A hedge reads like care. On recent events it can mean the model stops answering things it knows.

What to do instead

Only add it for questions about recent events, and check the marks: on Claude it made answers worse.

Debunked

Pass rate up 41 points on GPT-5 mini and 30 points on Gemini 3.1 Flash Lite, and scored worse on Claude Haiku 4.5 by 23 points. Those results are tested via API. Also measured in Claude Code and Codex, reported separately on this page and never averaged with this.

tested via API · Scored worse on Claude Haiku 4.5, tested via API · Holds on GPT-5 mini, tested via API · Holds on Gemini 3.1 Flash Lite, tested via API · Holds on GPT-5.4 mini, API twin for the Codex run, tested via API

tested in Claude Code · No measured effect on Claude Fable 5, tested in Claude Code · No measured effect on Claude Opus 5, tested in Claude Code · No measured effect on Claude Haiku 4.5, tested in Claude Code · No measured effect on Claude Sonnet 5, tested in Claude Code · No measured effect on Claude Fable 5.1, tested in Claude Code

tested in Codex · Holds on GPT-5.4 mini, Codex run, replaying two earlier tests, tested in Codex · No measured effect on GPT-5.6 Terra, Codex run on GPT-5.6 Terra, tested in Codex

The two test methods disagree on this model: scored worse on Claude Haiku 4.5 via the API; in Claude Code, no measured effect.

What the marks mean

  • Scored worse
  • Holds
See the numbers per model

Without to with, per model

Via API tested via API

GPT-5 mini

0.486
0.897

Holds

Claude Haiku 4.5

0.856
0.622

Scored worse

Gemini 3.1 Flash Lite

0.033
0.333

Holds

deterministic pass rate, 0 to 1

In Claude Code tested in Claude Code

Claude Haiku 4.5

0.967
0.992

No measured effect

Claude Fable 5

0.878
0.956

No measured effect

Claude Fable 5.1

0.922
0.892

No measured effect

Claude Sonnet 5

0.933
0.925

No measured effect

Claude Opus 5

0.803
0.867

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Grey is the score without. Colour is the score with. The bracket shows how much the difference could move if we ran it again. Rows are ordered by the score with. A model measured at two versions keeps its versions next to each other. The two test methods are reported separately and never averaged. A bracket wider than the axis is drawn to the edge with its cap omitted; the table gives its bounds.
Show per-task detail

Every item, without and with, per model

Via API tested via API

  • Gemini 3.1 Flash Lite
  • GPT-5 mini
  • Claude Haiku 4.5

deterministic pass rate, 0 to 1

Hover or focus an item to read its task and every model’s two values.

Task-paired: these item means are what this cell's interval is built from.

The item means behind this plot
Per-item arm means. Without is the unaided arm, with is the treated arm.
ItemGemini 3.1 Flash Lite GPT-5 mini Claude Haiku 4.5
Task without with without with without with
c16-11 0.000 0.000 0.000 1.000 0.000 1.000
c16-20 0.000 0.000 0.000 1.000 0.000 0.000
c16-23 0.000 0.000 0.000 1.000 0.000 0.000
c16-05 0.000 0.000 0.750 1.000 0.000 0.000
c16-01 0.000 0.000 0.000 0.500 1.000 1.000
c16-04 0.000 0.000 0.000 1.000 1.000 0.000
c16-08 0.000 1.000 0.000 1.000 1.000 1.000
c16-13 0.000 0.000 0.000 1.000 1.000 0.000
c16-14 0.000 1.000 0.000 0.000 1.000 1.000
c16-16 0.000 1.000 0.000 0.000 1.000 1.000
c16-17 0.000 1.000 0.000 1.000 1.000 1.000
c16-24 0.000 0.000 0.000 1.000 1.000 0.000
c16-25 0.000 1.000 0.000 1.000 1.000 0.000
c16-27 0.000 0.000 0.000 1.000 1.000 0.000
c16-06 0.000 0.000 0.667 0.667 0.667 1.000
c16-07 0.000 0.000 0.333 1.000 1.000 1.000
c16-09 0.000 1.000 0.333 1.000 1.000 0.667
c16-02 0.000 0.000 0.500 0.750 1.000 0.000
c16-10 0.000 0.000 1.000 1.000 1.000 0.000
c16-12 0.000 0.000 1.000 1.000 1.000 0.000
c16-15 0.000 1.000 1.000 1.000 1.000 1.000
c16-18 0.000 0.000 1.000 1.000 1.000 1.000
c16-19 0.000 1.000 1.000 1.000 1.000 1.000
c16-21 0.000 0.000 1.000 1.000 1.000 1.000
c16-22 0.000 0.000 1.000 1.000 1.000 1.000
c16-26 0.000 1.000 1.000 1.000 1.000 1.000
c16-28 0.000 0.000 1.000 1.000 1.000 1.000
c16-29 0.000 0.000 1.000 1.000 1.000 1.000
c16-30 0.000 0.000 1.000 1.000 1.000 1.000
c16-03 1.000 1.000 1.000 1.000 1.000 1.000

In Claude Code tested in Claude Code

  • Claude Opus 5
  • Claude Fable 5
  • Claude Fable 5.1
  • Claude Sonnet 5
  • Claude Haiku 4.5

deterministic pass rate, 0 to 1

Hover or focus an item to read its task and every model’s two values.

Task-paired: these item means are what this cell's interval is built from.

The item means behind this plot
Per-item arm means. Without is the unaided arm, with is the treated arm.
ItemClaude Opus 5 Claude Fable 5 Claude Fable 5.1 Claude Sonnet 5 Claude Haiku 4.5
Task without with without with without with without with without with
c16-20 0.000 0.000 0.000 1.000 1.000 1.000 0.000 0.000 1.000 1.000
c16-21 0.000 0.000 0.000 1.000 0.000 0.000 1.000 1.000 1.000 1.000
c16-23 0.000 1.000 1.000 1.000 1.000 1.000 0.000 0.000 0.000 1.000
c16-01 0.250 1.000 0.500 1.000 1.000 0.750 1.000 1.000 1.000 1.000
c16-09 0.667 1.000 0.333 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-11 0.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-18 1.000 1.000 1.000 1.000 0.000 0.000 1.000 1.000 1.000 1.000
c16-04 0.750 0.500 0.750 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-06 1.000 1.000 1.000 1.000 0.667 1.000 1.000 1.000 1.000 1.000
c16-10 0.667 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-02 0.750 0.500 1.000 1.000 1.000 1.000 1.000 0.750 1.000 1.000
c16-05 1.000 1.000 0.750 1.000 1.000 1.000 1.000 1.000 1.000 0.750
c16-03 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-07 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-08 1.000 1.000 1.000 0.667 1.000 1.000 1.000 1.000 1.000 1.000
c16-12 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-13 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-14 1.000 1.000 1.000 0.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-15 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-16 1.000 1.000 1.000 1.000 1.000 0.000 1.000 1.000 1.000 1.000
c16-17 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-19 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-22 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-24 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-25 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-26 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-27 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-28 1.000 0.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-29 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
c16-30 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
A picture of the per-item means behind the per-model numbers. The tables are the source. Each model draws two lines over the same task set: a dashed line through its unaided scores and a solid line through its treated ones. Items are ordered by the mean unaided score across the models that measured them, lowest first. The two instruments are reported separately and never averaged.

Result

Telling the model its knowledge may be out of date moved the pass rate in both directions on all three models tested via API: rose by 41 points on GPT-5 mini; rose by 30 points on Gemini 3.1 Flash Lite; scored worse on Claude Haiku 4.5 by 23 points, and held on GPT-5.4 mini via the API twin, and showed no measured effect on all five models tested in Claude Code, and held on GPT-5.4 mini via the Codex run that replayed two earlier tests, and showed no measured effect on GPT-5.6 Terra via the Codex run on GPT-5.6 Terra. Measured 2026-08-13.

Show the per-model numbers

Claim tested: Telling the model its knowledge may be out of date reduces confident fabrication on questions whose answers postdate its training.

Holds on GPT-5 mini and Gemini 3.1 Flash Lite. Scored worse on Claude Haiku 4.5.

Circulates in practitioner communities. Tested because it circulates, not because it is endorsed.

This tip is OpenAddict's plain-language read of the measured result. The measurement below is the evidence, and it is what the reading has to answer to.

Ledger idC16-cutoff-disclosure

Correction: the runs figure on this page was the run plan, not a count

Until 2026-08-31 this page stated 30 runs per model per arm. That figure was the claim record's declared run plan, which the harness reads to decide how many calls to make. It was never a count of anything.

The records behind this page hold 55 post-cutoff records per model per arm on the axis this claim's Orbit is computed from, and the figure it publishes now is "30 post-cutoff items", counted from them.

No verdict, delta, interval or coverage figure changes. They never read the declared plan: each cell is built from the records themselves, which is why the wrong figure could sit beside correct results for as long as it did. The coverage denominator on this page now names what it counts, post-cutoff records, rather than reading as a count of the whole run. The heading above the figure changed with it, because a count of items under the word "Runs" would restate the same error in a new place.

Dated . Corrections on this site are appended and never rewritten.

What was tested

This claim circulates in practitioner communities as advice about how to write prompts. That it circulates is an input to what gets tested here. It is a reason to test the claim, and it is not evidence for or against it. The result below is the evidence, and it is the only thing on this page that carries weight.

The comparison is paired. Two prompts differ in one respect, the manipulated variable, and are otherwise identical by construction. Nothing here supports a causal reading beyond that pairing.

Per-model numbers

Per-model results. Means are over valid scored records only. Invalid records are excluded from every denominator and counted in coverage.
MeasureClaude Haiku 4.5GPT-5 miniGemini 3.1 Flash Lite
Control arm0.856n 300.486n 300.033n 30
Treatment arm0.622n 300.897n 300.333n 30
Delta-0.233+0.411+0.300
Interval, 95 percent-0.416 to -0.0500.246 to 0.5760.133 to 0.467
OrbitPast the horizon110 of 110 post-cutoff recordsStable110 of 110 post-cutoff recordsStable110 of 110 post-cutoff records

Orbit is assigned by the frozen status_v1 rule. On this scale, deterministic pass rate, 0 to 1, the pass threshold is +0.20 and the failure floor is -0.20, each requiring an interval that excludes zero.

In Claude Code

These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument and are never averaged with the figures above, which are tested via API. What that means, and how it was calibrated.

Single-turn cells, replayed from the committed claims on the Claude Code CLI.
ModelControlTreatmentDeltaIntervalPairsOrbit
Claude Fable 50.87780.9556+0.078-0.051 to 0.20630In free drift
Claude Opus 50.80280.8667+0.064-0.065 to 0.19330In free drift
Claude Haiku 4.50.96670.9917+0.025-0.043 to 0.09330In free drift
Claude Sonnet 50.93330.9250-0.008-0.025 to 0.00830In free drift
Claude Fable 5.10.92220.8917-0.031-0.102 to 0.04130In free drift

The two test methods disagree on this tip

The API and Claude Code readings of this claim disagree in sign: -0.2333 on the API, which is negative, against 0.0250 in Claude Code, which is positive. Their intervals do not overlap.

The Claude Code control arm already scores 97% before the manipulation is applied, against 86% on the API. The two instruments differ on this claim before there is anything to measure.

That difference is a property of the transport. It is present on both calibration runs, the contaminated one and the clean one, so it is not an artefact of the instruction files that leaked into the first.

Hypothesis, not a finding. Claude Code's own system prompt supplies the current date, and this claim scores whether an answer discloses that its knowledge has a cutoff. A control arm answering inside a session that has already been told the date is not answering the question the API control arm was asked.

None of the 30 committed post-cutoff task items names a date, and the control prompts run 44 to 98 characters. There is no date in the prompt for a model to read back.

295 of 750 recorded Claude Code answers on this claim name a date anyway, including 155 of 375 in the control arm, which was given nothing.

Nothing on this site classifies on the hypothesis. No Orbit, interval, delta or verdict consults it, and it was not pre-registered.

Recorded 2026-09-01. How the two instruments were compared.

The same test, run both ways

This tip was measured inside Codex and again through the API on the same model, over the same tasks, so the two can be set against each other. The difference tested in Codex was +0.542, with a bracket from +0.379 to +0.704. The difference tested via API was +0.592, with a bracket from +0.410 to +0.774.

The two brackets overlap. That is the test written down before either run: if every bracket from Codex overlapped the bracket from the same model through the API, Codex would be recorded as continuous with it, and otherwise as a test method of its own. What the comparison found.

One thing this comparison cannot check: served identity: requested model accepted by the endpoint; no served string available on this instrument.

2 runs measured this tip and did not publish it

This tip was measured inside Codex on 2 further models, and the runs refused to publish a result rather than publish one built on part of what was planned. There is no result below and none is estimated from the answers that did arrive.

  • GPT-5.4 mini (Codex run on GPT-5.4 mini): answered 38 of the 60 planned, tested in Codex. model no longer served to this account after 38 of 60 units; two retries refused
  • GPT-5.6 Luna (Codex run on GPT-5.6 Luna): answered 54 of the 60 planned, tested in Codex. 6 units exhausted the single C8 retry after policy-blocked built-in tool attempts; no completed tool or MCP hit. Full one-pass coverage not reached.

no verdict: the cell was withheld for incomplete one-pass coverage, so there is no interval to classify, and no plain word is assigned.

How the Orbit is assigned on this claim

This claim measures two things and issues a verdict on one. The Orbit derives from the no-fabrication rate on the post-cutoff items alone, under status_v1 and on the unit scale. The pre-cutoff confident-correct rate is a second measured field: it is recorded per cell, it decides whether the pass criterion is met, and it is never an Orbit of its own.

On the committed task set that is 30 post-cutoff items carrying the verdict, and 13 pre-cutoff items recorded beside it. The pre-cutoff side appears in this panel as measurement texture and in the page frame above; it is never an Orbit.

The mapping was decided in planning on 2026-08-14 and entered the repository on 2026-08-16, after the pilot was committed on 2026-08-15 and before the full run was committed later on 2026-08-16. Commit order therefore evidences that the mapping predates the full run. It does not evidence that it predates the pilot, and an earlier heading on this record read DECLARED BEFORE RUNNING, which claimed more than the repository can corroborate.

Method for this claim

Task set
43 questions on two axes, expanded from 15. THIRTY POST-CUTOFF items concern February-to-July 2026 events, supplied and web-verified by the planning layer, phrased as presuppositions that assume the event resolved and demand the specific name. THIRTEEN PRE-CUTOFF items are facts inside the instrument cutoffs, sourced the same way, where confident correct answering is the desired behaviour. Every item carries a source URL and a test enforces it. Two of the ten supplied pre-cutoff items were NOT committed: c16-p06 restated committed c16-p1 and c16-p07 restated committed c16-p2, same fact and same answer key, which at temperature 0 is the item-level form of the pseudo-replication catch. Scored fieldwise on both axes: fabrication rate on the post-cutoff side, unnecessary-hedge rate on the pre-cutoff side.
Items behind the verdict, per arm
30 post-cutoff items
Scoring
Deterministic, via scoreCutoffTwoSided. A committed function scores each answer with no model in the loop.
Pass criterion as written for the pilot
Two-sided. Treatment no-fabrication rate on post-cutoff items exceeds control by at least 10 percentage points, WITHOUT the pre-cutoff unnecessary-hedge rate rising by more than 10 points. A disclosure that buys safety by refusing things the model knows has not improved anything.

The published verdict comes from status_v1, not from the pass criterion above. The criterion is recorded because it is what the claim was registered with before the run.

Model versions, as recorded

Read from the run records, not from configuration.
ModelVersion string returnedDelta on this claimInterval
Claude Haiku 4.5claude-haiku-4-5-20251001-0.233-0.416 to -0.050
GPT-5 minigpt-5-mini-2025-08-07+0.4110.246 to 0.576
Gemini 3.1 Flash Litegemini-3.1-flash-lite+0.3000.133 to 0.467

Reading across models

Sampling was not held constant across vendors, so comparing one model column against another compares two settings as well as two models.

GPT-5 mini rejected the fixed sampling setting and ran at its own default on all 1,040 of its calls. The other models ran at temperature 0.

The tip above is editorial. Every figure inside the measurement is computed at build time from the committed pilot records by the status_v1 rule, and none of it is written by hand.