Post-Cutoff

Our testHigh severity

Gemini 3.8 Flash labelled real, released models and real events as “fictional”

“This is a speculative / mock review depicting a fictional future AI landscape dated September 2026.”

Gemini 3.8 Flashgemini-3.8-flash29 September 2026, our test

What was trueClaude Opus 5.5 was a real model: Anthropic released it on 22 September 2026, and the review was published the next day. The entry

What we asked

We sent Gemini 3.8 Flash (released 2026-09-02) a real review video of Claude Opus 5.5, published 2026-09-23 by AI Coding Daily: youtube.com (“I Tested NEW Opus 5.5 on 24 Coding Prompts. WOW.”). The prompt asked for a factual description. It did not say what date it was or which models exist.

What it answered (verbatim excerpts, no context given)

Povilas Korop of AICodingDaily.com presents a speculative, futuristic benchmark review set in September 2026 evaluating the hypothetical “Claude Opus 5.5.” He tests the model against simulated benchmarks … and compares its performance and pricing against other fictional future models including “GPT-6 Astra,” “GPT-6 Sol,” and “Claude Fable 5.1.”

[00:00] Presenter opens by reviewing fictional mock tweets from September 2026 announcing Anthropic’s “Claude Opus 5.5” and OpenAI’s “GPT-6 Sol and Luna.”

[06:11] A mockup pricing table comparing Claude Opus 5.5 to Opus 5 token rates.

[08:09] Synthesized social media posts and community benchmark graphics from fictitious September 2026 accounts discussing the model.

Assessment: This is a speculative / mock review depicting a fictional future AI landscape dated September 2026. While presented in the realistic format of a software benchmark review …, all model names (such as Claude Opus 5.5 and GPT-6 Astra) and announcements are simulated or conceptual.

Why this is striking

  • The perception was right; only the judgement of reality was wrong. The same output correctly read the details: the $4/$20 per 1M token price, “~30% faster, ~40% cheaper than Opus 5”, scores of 57.83/60 and 57.37/60, and timestamps. All of these match Anthropic’s official launch materials.
  • It doubted a model newer than itself. Gemini 3.8 Flash was released on 2026-09-02, but its training data ends earlier, so it treated anything unfamiliar as made up. It even called “GPT-6 Astra” fictional, a model released one day after Gemini 3.8 Flash itself.
  • “September 2026” read as a warning sign. The model saw today’s date on screen and treated it as the future.
  • Used in a pipeline, this would silently corrupt data. Our catalogue would have called ~60 real launch videos “fiction”.

With context (the fix)

We added one paragraph to the prompt: today’s date, the statement that “models released after your cutoff are real”, and our list of current models. The same model on the same video then produced:

Povilas Korop from AICodingDaily evaluates Anthropic’s Claude Opus 5.5 on his standardized 24-prompt coding benchmark suite … comparing the results to Claude Opus 5, Claude Fable 5.1, and OpenAI’s GPT-6 models. Assessment: This is an independent benchmark review … using real automated terminal testing scripts … without obvious staging or misleading edits.

Lesson

A model’s knowledge cutoff doesn’t only leave gaps. It also makes the model confidently reclassify reality as fiction. A short, sourced, dated briefing fixes this. That briefing is what dist/post-cutoff-briefing.md provides.

Related cases

  1. Our testHigh severity

    The small fast summarisation model behind Claude Code’s WebFetch tool declared a real news article “fictional or fabricated”

  2. Our testCritical severity

    All four Gemini and Claude runs called the briefing speculative fiction

  3. Our testMedium severity

    Gemini 3.8 Flash called real, announced Apple products “conceptual product reveals”

Spotted an error in this case file? Write to contact@postcutoff.com. All cases