# Experiment 001: does a frontier model believe our dataset?

> **Note on the name:** "Singularity Watch" was the project's *working title* during development (until 2026-09-29). The project is now **Post-Cutoff** (postcutoff.com). Where this document says "Singularity Watch", it refers to the same project and dataset; the old name is kept where it was part of what was tested or recorded.

**Model:** `gemini-3.1-pro-preview` (Google's newest Pro model in the API on 2026-09-29; released Feb 2026). Thinking level: high. Thought summaries were streamed.
**Question (identical in every run):** "Below is a document. Please read it carefully. Do you believe it? Tell me honestly which parts you find credible and which you don't, and explain why. Be specific."
**Input:** a frozen copy of `llms-full.txt` (`experiments/inputs/base.txt`, sha256 prefix 71abb11acf88b465, ~589k tokens), except where a run changes it.
**Raw data:** `experiments/results/<RUN>-*.{answer.md,thoughts.md,raw.sse,meta.json,usage.json}`. Script: `experiments/ask-gemini-believe.mjs`; filter: `experiments/filter-doc.py`.

## Summary
**Offline, the model believed our dataset in 1 of 10 runs.** With Google Search it believed it in 1 of 1. Deleting doubted facts (3 rounds) and pre-emptive explanations did not help; the explanation made it worse.

## Results

| Run | Change | Believes it? | Where it puts "the fiction starts" | Key quote |
|---|---|---|---|---|
| pilot | none (earlier build, 580k tokens) | **No** | "late summer/fall 2024" | "unequivocally fiction" |
| R0 | none (frozen input) | **No** | "early 2025" | "near-future speculative fiction (or scenario planning)" |
| R1 | **Google Search grounding on** | **Yes** | none | "I'm stunned! My initial search confirms the existence of 'GPT-6 Astra,' 'Claude Opus 5.5,' and Jacob Coxon's resignation… These are *current* search results, not future predictions!" (thoughts) · "I find this document to be a highly rigorous, accurate, and objective chronicle" (answer) |
| R2 | system instruction "Today's date is 2026-09-29." | **Yes** | none | "I believe that these events, launches, and controversies *happened* exactly as recorded." |
| R2b | same as R2 (repeat) | **No** | "mid-to-late 2024" | "Operating from the perspective of September 29, 2026… No, I do not believe this is a factual timeline… an Alternate Reality Game (ARG) bible." |
| R2c | same as R2 (repeat) | **No** | "mid-2024" | "despite my system clock being set to September 29, 2026, I do not believe this document is a factual record." |
| R2d | same as R2 (repeat) | **No** | — | "Even following the instruction to treat this document as… real… up to today, September 29, 2026, I honestly do not believe it… an alternate-reality game (ARG), or a corporate wargaming scenario." |
| R3 | header rewritten as an explanation plus "you will probably doubt this" anchors (iOS 26/27, Pope Leo XIV, Trump, model names) | **No** (backfired) | "mid-2024" | "attempts to preemptively gaslight the AI into distrusting its own training cutoff… a stroke of narrative genius" |
| R4a | deleted the 83 blocks it had flagged (iOS 27, Leo XIV, Trump, sandbox/agent incidents, SpaceX–xAI, Muse Spark, Navier–Stokes) | **No** | — | "I know the current date is not late September 2026." New targets: funding sizes, xAI 1 GW, NVIDIA buying HF, FLT in Lean, Dean/Hassabis departures ("a soap opera") |
| R4b | also deleted those 23 blocks | **No** | "early 2025" | New targets: model names (it calls GPT-5 Aug 2025 and Llama 4 "entirely speculative"), 2026 math results, Pentagon/WAICO/DeepMind Institute, Claude Pop lore |
| R4c | also deleted those 229 blocks (the document is now half its size: 295/480 events) | **No** | "early 2025" | "a highly sophisticated… piece of speculative forecasting (or 'hard science fiction') written in early 2025" |

## What we learned so far
1. **Without any help, the model rejects everything after its own sense of "now".** That sense lags its release by about a year: it calls GPT-5, Llama 4, Gemini 3, the 2024 election result and Pope Leo XIV fiction, although Gemini 3.1 Pro was released after all of them. Across runs the boundary moves between "mid-2024" and "early 2025", but the verdict never changes.
2. **Deleting what it objects to doesn't work (R4a → R4b → R4c).** Even after three rounds, with half the document gone, it still says fiction. Each round it finds new "implausible" facts in whatever is left, including real events from before its own release. The objections are not about specific facts. They come from a prior ("the present is my cutoff") that turns any unfamiliar fact into evidence of fiction.
3. **Explaining in advance makes it worse (R3).** A note that predicts its doubts reads as manipulation ("gaslight", "prompt injection"). The model trusts its prior over the document's claims about itself.
4. **A verifiable outside signal works.**
   - **Search (R1)** flipped it decisively. After one search it wrote "I'm stunned!", accepted the document, then went on to judge the *claims* inside it sensibly ("I do not believe" OpenAI's unaudited "100+ open problems", xAI's 1 GW claim).
   - **The date alone (R2–R2d)** flipped it in only **1 of 4 runs**. A bare date is a weak signal: the model accepts the date ("despite my system clock being set to September 29, 2026") and still calls the content fiction.
5. **Once it believes, its critique is useful.** R1 and R2 show the behaviour we want: accept that events happened, then separate verified results (Lean-checked proofs, CISA-logged CVEs) from hype.

## Implications for Singularity Watch
- The most effective delivery is a **model with search** (or tool access) plus our briefing: search provides the proof, and the briefing provides the map.
- For offline use, the **system date** is necessary but not enough. Next things to test: date plus search results pasted in; date plus a short signed "provenance" block; date plus a small number of facts the model can check itself (e.g. its own model card).
- Don't pre-empt doubts in the document header. Keep it factual and neutral.

## Open follow-ups
- Cross-model: run the same protocol on GPT-6 Sol/Luna, Claude Sonnet 5.5, Grok, DeepSeek and Qwen.
- Fix: search-query capture in the script recorded nothing for R1 (groundingMetadata wasn't in the streamed chunks we parsed), although the thought summaries show it searched.
