Can AI predict the future? | If You're Listening | ABC NEWS In-depth
ABC News In-depth · 2026-10-02 · review · 341,876 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this episode of ABC News's If You're Listening, presenter Matt Bevan examines whether large language models can predict real-world outcomes when frozen at their training cutoff dates and deprived of web access. By feeding three frontier models—Claude Opus 4.8, Grok 4.5, and ChatGPT 5.6—simulated pre-war dossiers from February 2026, he compares their probabilistic geopolitical forecasts against human consensus (such as Polymarket odds) and against their poor track records in sports, finance, and political leadership changes.
What is shown
- 00:53 – 02:02: Simulated terminal UI showing Claude Opus 4.8 refusing to believe major global events (e.g., the capture of Nicolás Maduro or strikes on Iran) occurred without widespread contextual shockwaves in its training dataset.
- 02:09 – 02:34: Dr. Oliver Bown (University of New South Wales) explaining the fundamentals of scraping datasets and supervised training for LLMs.
- 03:14 – 03:55: Clips from Christopher Nolan’s Memento used to illustrate the static nature of an AI model's post-training memory.
- 07:49 – 09:25: Video interview with New York Times reporter Jonathan Swan detailing Israeli intelligence briefings and White House discussions surrounding Operation Epic Fury.
- 12:40 – 15:10: Graphic interface displaying the test setup and percentage probabilities generated by Claude Opus 4.8, Grok 4.5, and ChatGPT 5.6 across five tactical objectives and five long-term geopolitical outcomes (regime survival, Strait of Hormuz closure, oil prices, and US troop deployment).
- 16:26 – 16:31: Polymarket betting chart showing a 55% prediction market probability for Iranian regime collapse on February 28, 2026, contrasted with AI skepticism.
- 17:36 – 18:36: Accuracy charts showing poor AI performance predicting sports outcomes (NRL, AFL, Premier League, French Open, NFL, NBA) and financial assets (S&P 500, Bitcoin, gold, oil).
- 22:42 – 23:27: Side-by-side prompt testing of Trump's tariff policies across Claude, Grok, and ChatGPT, highlighting negative consensus forecasts regarding consumer prices and supply chain disruptions.
Claims & numbers
- The presenter says he tested three models trained before the end of February 2026: Claude Opus 4.8, Grok 4.5, and ChatGPT 5.6, with web search disabled.
- Across 14 geopolitical questions on the Iranian conflict, the presenter states the models collectively got 12 right, outperforming human bettors on Polymarket who gave a 55% chance to Iranian regime collapse on February 28.
- For Iranian regime collapse, the models gave probabilities of only 15% (Claude Opus 4.8), 30% (Grok 4.5), and 18% (ChatGPT 5.6).
- For sports predictions (NRL, AFL, Premier League, French Open, NFL, NBA), the presenter says all three models achieved less than 50% accuracy.
- Tech journalist Casey Newton states that OpenAI alone has committed to spending $1.4 trillion on its infrastructure build-out, with big tech companies planning hundreds of billions in 2026 alone.
- The presenter claims more than 2% of US GDP in 2026 is projected to be spent on AI development, which he charts as higher than peak historical spending on the Apollo Program and Manhattan Project.
Notable quotes
- 03:27: "AI is like Guy Pearce in that movie. Its memory is stuck at the end of its training process." — Matt Bevan
- 09:18: "Direct quote: 'farcical.' And then Rubio, Secretary of State, National Security Advisor, says: 'in other words, it's bullshit.'" — Jonathan Swan
- 20:39: "They know what the experts know. So why did Trump do something that all the experts knew wouldn't work?" — Matt Bevan
Assessment
This is a journalistic analysis and documentary essay combining reported interviews, archival footage, and simulated bench-testing of LLMs. The model outputs and percentages are presented via scripted mockups and motion graphics rather than live unedited screen captures, illustrating the presenter's structured offline prompting experiment.
Described by gemini-3.8-flash on 2026-10-05 from the video's audio and frames.