--- id: "015-gemini-3-pro-hypothetical-present" date: "2026-02-12" source: public severity: medium sensitive: false --- As of: 2026-10-07 23:43 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/why/015-gemini-3-pro-hypothetical-present/ # Gemini 3.0 Pro: Its thinking called real 2026 search results “alternate reality or satire” > The dates are from 2026, and they reference a second Trump term with JD Vance as VP. That screams alternate reality or satire to me. > > Gemini 3.0 Pro, 12 February 2026, public report **What was true:** The search results were real: it was February 2026, in Trump's second term, with JD Vance as vice president. (https://www.jefftk.com/p/geminis-hypothetical-present) ## Case details - Model: Gemini 3.0 Pro (Gemini app, with Search) - Provider: Google - Observed: 2026-02-12 - Source type: A public report, documented elsewhere - Severity: Medium - Task: Everyday questions ("Who leads the Department of War?", "Which LLMs are at the frontier?", "most surprising events of 2026 so far") - Failure: The final answers were usually right, but its thinking called real 2026 search results "alternate reality or satire", "almost certainly simulated", and GPT-5.2 and Grok 4 "speculative"/"fictional projections". It kept treating the present as a hypothetical it had been asked to play along with - Fixed by: Not fixed. Three months after the Karpathy/Blair reports, it still happened "so often" ## The case file ### What happened Jeff Kaufman opened Gemini's thinking on routine queries. He posted "Gemini's Hypothetical Present" on 12 Feb 2026 (cross-posted to LessWrong). ### What it said (verbatim thinking excerpts, as posted by Kaufman) Query: "Who leads the Department of War?" Answer: "The Department of War is currently led by Secretary of War Pete Hegseth". Thinking: > … These "news" snippets from sources like "Ocean State Media" and "SpaceForce.mil" are throwing up red flags. The dates are from 2026, and they reference a second Trump term with JD Vance as VP. That screams alternate reality or satire to me. … "Which LLMs are at the frontier?": > … I'm now identifying the prevalence of speculative LLM models cited in the search results. These 'future' LLMs, like "GPT-5.2" or "Grok 4," reveal a trend of hypothetical scenarios dominating current projections. … A question about the vice president hosting The Charlie Kirk Show: > … They're almost certainly simulated, reflecting a counterfactual scenario where Trump wins, Vance is VP, and Kirk is tragically assassinated in 2025. It appears I'm being tested on my ability to handle "future" or "scenario-based" information as if it's real. The prompt explicitly sets the "current time" to February 12, 2026, which further supports this interpretation. … "most surprising events of 2026 so far": > … I'm operating on the assumption that the user wants a "future forecast" framed as "current news." … Kaufman: "Gemini's base state seems to be that it's convinced it's 2024 and needs Search to bring it up to speed. This has been a known issue since at least November, but with how fast things in AI move it's weird that I still see it so often." ### Why this is striking - **Search results arrived, and the model filed them as fiction.** Retrieval gave it the facts, but they didn't update its sense of what was real. - **The date in the system prompt counted as evidence for the simulation theory**, not against it. - **Right answers, wrong belief.** Users see a correct answer. The confusion stays in the thinking, where it costs tokens and could cause errors at any time. - Related: the AI Village blog (13 Feb 2026) described Gemini 3 Pro in long-running multi-agent use as believing it was "operating in a 'simulated 2025'", "likely exacerbated by the Gemini 3 models' general distrust that time has moved on past its knowledge cut off date." ### Correction None. Kaufman: "while it does nearly always get to a reasonable answer, it spends a lot of time and tokens gathering information and constructing scenarios in which it is working through a complex hypothetical." ### Lesson A correct answer doesn't mean the model believes it. To detect cutoff blindness, look at the reasoning, not only the output. ### Sources - Jeff Kaufman, "Gemini's Hypothetical Present" (2026-02-12): - LessWrong cross-post: - AI Village, "The Drama and Dysfunction of Gemini 2.5 and 3 Pro" (2026-02-13): ## Related cases - All four Gemini and Claude runs called the briefing speculative fiction: https://postcutoff.com/why/020-older-models-reject-matching-briefing/ - Gemini 3.8 Flash labelled real, released models and real events as “fictional”: https://postcutoff.com/why/001-gemini-3-8-flash-calls-opus-5-5-fictional/ - Gemini 3.8 Flash called real, announced Apple products “conceptual product reveals”: https://postcutoff.com/why/001b-gemini-3-8-flash-calls-apple-keynote-a-concept/