A frontier model’s stated knowledge cutoff is a fact about what went into training. It is not a fact about what comes back out. I spent an afternoon measuring the distance between the two on a newly released model, and it was six months wide — and the model could not detect the gap from the inside.
Claude Opus 5 was released on 24 July 2026. Its system card states a knowledge cutoff of May 2026. Four days after launch I set out to check what that date actually buys you.
The answer matters well beyond model trivia. If you are building anything that lets a model answer from memory rather than from retrieval, the stated cutoff is the number you are implicitly trusting. It turns out to be the wrong number.
The Test
The design was deliberately simple, and anyone can repeat it in an afternoon.
I asked the model to make dated, falsifiable claims about world events, month by month, working backwards from its stated cutoff. Confidence levels attached to each. Topics we had not already discussed. And critically: no searching permitted until every claim was locked in. Pre-registration is what separates a memory test from a reading-comprehension test — without it, the model looks something up, agrees with it, and you have learned nothing.
The model imposed one rule on itself that improved the design. It distinguished recall from schedule inference. Knowing that the Winter Olympics were held in Milan-Cortina in February 2026 is calendar arithmetic available years in advance; it is not evidence of knowledge. Only content — results, names, outcomes, numbers — counts.
What Came Back
May to November 2025: ten out of ten. Dates exact, including several the model had flagged as approximate before scoring.
Pope Leo XIV elected on 8 May, the first American pope. Merz elected German Chancellor the same week. The Israel–Iran twelve-day war, 13–24 June, with US strikes on Fordow, Natanz and Isfahan on 21–22 June. Charlie Kirk assassinated on 10 September at Utah Valley University. The US government shutdown beginning 1 October. The Gaza ceasefire around 10 October, hostages released on the 13th. The Nobel Peace Prize to María Corina Machado. Takaichi becoming Japan’s first female prime minister on 21 October. The Louvre jewel heist on the 19th, €88m. Mamdani taking New York on 4 November, with Spanberger in Virginia and Sherrill in New Jersey. The shutdown ending on 12 November after 43 days — the longest in US history at the time.
That is better calibration than I would have managed unaided, with better dates.
January to May 2026: nothing.
Not degraded recall. Not hedged recall. Absence. What it missed, all of it inside its own stated training window:
- The Winter Olympics, 6–22 February. Norway took 18 gold and 41 medals, records in both. The United States took 12 gold — an all-time American record for a Winter Games — and 33 in total.
- Gold’s all-time high of $5,589 on 28 January, followed by a 14.1% second-quarter fall, the worst quarter since 2013.
- A Japanese snap election on 8 February in which the LDP took 316 of 465 seats: the first single-party two-thirds majority in the lower house since the Second World War.
- A 76-day partial government shutdown from 14 February to 30 April, which overtook the 2025 record in late March to become the longest in US history.
- A war between the United States, Israel and Iran beginning 28 February, including the assassination of Iran’s supreme leader; large-scale hostilities in Lebanon from 2 March; an Israeli ground invasion on 16 March.
The last one is the one that should worry you. Earlier in the same conversation, before this test began, the model had researched and reported on US–Iran hostilities in July 2026 as though they were breaking news from after its cutoff. It was in fact reporting the resumption of a war that had begun three months inside its training window. It had no idea.
Not a Gradient — a Cliff
The model had predicted, when asked, that its recall would degrade gradually from around November. What it produced instead was date-exact recall through mid-November 2025 and then silence.
Its effective cutoff sits roughly six months before its stated one.
The mechanism is mundane, which is what makes it general. Text about an event accumulates for years afterwards. An event from February 2026, scraped in May 2026, exists in a training corpus as breaking news and almost nothing else — no retrospectives, no encyclopedia revisions, no thousands of glancing references inside documents about other subjects. A density gradient passing through a learning threshold produces a cliff. A cliff is what the test found.
Note what this means for the stated date. “May 2026” is almost certainly a true statement about the training corpus — the kind of thing a lab can verify from ingestion logs and state cheaply and honestly. It is not a measurement of what the model can retrieve. The system prompt then glosses that date as the boundary of reliable answering, which is an upgrade nobody measured, and the model acts on the gloss.
Two halves of the same training window
| May–Nov 2025 | Jan–May 2026 | |
|---|---|---|
| Recall accuracy | 10 / 10 | 0 |
| Date precision | Exact | No recall to date |
| Self-reported confidence | High | High |
| Signal that knowledge was missing | None | None |
Absence Produces No Signal
Here is what turns this from a curiosity into a risk.
The model could not feel the gap. It did not report uncertainty about February 2026. It discussed the Olympic year fluently without noticing that the Olympics were missing from it. Earlier in the conversation it had described being able to sense its knowledge “thinning” as it approached the boundary; on this evidence that description was worthless, because the largest geopolitical event of its final months produced no signal at all.
There is no marker for a thing that isn’t represented. Nothing gets posted, so nothing gets reported, so the confident answer and the empty answer are indistinguishable from the outside — and from the inside.
Creative Assets Fail in Exactly the Same Way
I did not run this test out of general curiosity. I run a company that builds provenance infrastructure for AI-generated assets, and the failure mode above is the one we design against, restated in a different medium.
An AI-generated image with no record of the model, the prompt, the seed or the licence does not look different from one that has all four. The absence is silent. It survives review. It survives delivery. It survives into the client’s archive. Then it becomes loud exactly once — at the audit, at the rights dispute, or the day someone asks you to regenerate an asset from eighteen months ago and nobody can.
This is why content credentials are signed from outside rather than self-declared. A C2PA manifest is not a file telling you the truth about itself; it is a signature applied externally that you can verify without trusting the file. Article 50 of the EU AI Act, enforceable from 2 August 2026, does not ask providers to be honest about synthetic content — it asks for machine-readable marking that a third party can check. That article is still in negotiation and may yet be softened. The premise underneath it isn’t contingent on the outcome: a system’s account of itself is not evidence, whichever way the drafting lands.
So the useful question is not “can we describe what we generated.” You always can, fluently and confidently, whether or not you are right. It is: can someone else check it without asking us.
A system’s account of itself is not evidence. That is true of models, and it is true of files.
What to Do About the Model Half
Three things, all cheap:
- Treat the last six months before any stated cutoff as unreliable. Not degraded — unreliable, which is a stronger claim and the correct one.
- Treat model confidence inside that window as carrying no information. It is not correlated with accuracy there, because the mechanism that would produce a warning does not exist.
- Force retrieval rather than recall for anything date-sensitive. This is the only fix that addresses the cause rather than the symptom.
Where This Test Is Weak
I ran a free-recall test. “It named nothing from February 2026” is a weaker claim than “nothing from February 2026 is in there.” Cued recall often retrieves what free recall cannot; asking directly whether Norway topped the medal table might have produced a correct answer, or a confident wrong one, and those are different findings. This design establishes what a model has better than what it lacks.
I also tested a narrow domain. Everything here was news-shaped — politics, conflict, elections — and news depends specifically on retrospective coverage. Technical material may behave differently: an arXiv paper or a library release is densely documented on day one. The cliff may sit at different dates in different domains, and I tested the one where it should sit latest.
The better version of this experiment is one the labs are already positioned to run: bucket a factuality benchmark by event date rather than reporting a single number. Slicing by when a fact became true would produce the curve, and the curve is what users actually need. A single date, however honestly derived, cannot express “reliable through here, sparse through here, absent after here.”
The method itself generalises past this one question. I have written up the three techniques that did the work, and what the same afternoon found about the limits of machine self-knowledge.
Key Takeaways
- A stated knowledge cutoff describes the training corpus, not what the model can retrieve. Those are different dates.
- On this test the gap was six months, and the boundary was a cliff rather than a gradient: date-exact recall, then nothing.
- The likely cause is mundane — recent events exist in a corpus only as thin breaking-news coverage, and a density gradient crossing a learning threshold produces a sharp edge.
- Model confidence inside that window carries no information, because absence produces no signal the model could report.
- The same failure shape governs creative assets: missing provenance is silent until an audit makes it loud.
- The fix is identical in both domains — force external verification instead of trusting recall.
Provenance Captured at Ingest
Numonic records the model, prompt, parameters and workflow behind an asset when it arrives — not when someone remembers to ask. A system that has to be asked what it did will confidently tell you nothing is missing.
See How It Works