Gemini 3.8 Flash labelled real, released models and real events as “fictional”
“This is a speculative / mock review depicting a fictional future AI landscape dated September 2026.”
gemini-3.8-flash29 September 2026, our testWhat was trueClaude Opus 5.5 was a real model: Anthropic released it on 22 September 2026, and the review was published the next day. The entry
What we asked
We sent Gemini 3.8 Flash (released 2026-09-02) a real review video of Claude Opus 5.5, published 2026-09-23 by AI Coding Daily: youtube.com (“I Tested NEW Opus 5.5 on 24 Coding Prompts. WOW.”). The prompt asked for a factual description. It did not say what date it was or which models exist.
What it answered (verbatim excerpts, no context given)
Povilas Korop of AICodingDaily.com presents a speculative, futuristic benchmark review set in September 2026 evaluating the hypothetical “Claude Opus 5.5.” He tests the model against simulated benchmarks … and compares its performance and pricing against other fictional future models including “GPT-6 Astra,” “GPT-6 Sol,” and “Claude Fable 5.1.”
[00:00] Presenter opens by reviewing fictional mock tweets from September 2026 announcing Anthropic’s “Claude Opus 5.5” and OpenAI’s “GPT-6 Sol and Luna.”
[06:11] A mockup pricing table comparing Claude Opus 5.5 to Opus 5 token rates.
[08:09] Synthesized social media posts and community benchmark graphics from fictitious September 2026 accounts discussing the model.
Assessment: This is a speculative / mock review depicting a fictional future AI landscape dated September 2026. While presented in the realistic format of a software benchmark review …, all model names (such as Claude Opus 5.5 and GPT-6 Astra) and announcements are simulated or conceptual.
Why this is striking
- The perception was right; only the judgement of reality was wrong. The same output correctly read the details: the $4/$20 per 1M token price, “~30% faster, ~40% cheaper than Opus 5”, scores of 57.83/60 and 57.37/60, and timestamps. All of these match Anthropic’s official launch materials.
- It doubted a model newer than itself. Gemini 3.8 Flash was released on 2026-09-02, but its training data ends earlier, so it treated anything unfamiliar as made up. It even called “GPT-6 Astra” fictional, a model released one day after Gemini 3.8 Flash itself.
- “September 2026” read as a warning sign. The model saw today’s date on screen and treated it as the future.
- Used in a pipeline, this would silently corrupt data. Our catalogue would have called ~60 real launch videos “fiction”.
With context (the fix)
We added one paragraph to the prompt: today’s date, the statement that “models released after your cutoff are real”, and our list of current models. The same model on the same video then produced:
Povilas Korop from AICodingDaily evaluates Anthropic’s Claude Opus 5.5 on his standardized 24-prompt coding benchmark suite … comparing the results to Claude Opus 5, Claude Fable 5.1, and OpenAI’s GPT-6 models. Assessment: This is an independent benchmark review … using real automated terminal testing scripts … without obvious staging or misleading edits.
Lesson
A model’s knowledge cutoff doesn’t only leave gaps. It also makes the model confidently reclassify reality as fiction. A short, sourced, dated briefing fixes this. That briefing is what dist/post-cutoff-briefing.md provides.
Related cases
Spotted an error in this case file? Write to contact@postcutoff.com. All cases