AI progress is speeding up dramatically - here's why you should care
Sky News · 2026-09-22 · review · 1,534,679 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Sky News technology correspondent Rowland Manthorpe presents an analytical studio report on why AI researchers are alarmed by the accelerating pace of artificial intelligence development. He examines benchmarks, unintended model behaviors, reliability discrepancies, and data from frontier labs showing that AI is transitioning from assisting researchers to autonomously leading technical tasks.
What is shown
- [00:04] – Studio screen showing an early math word problem from the AddSub benchmark ("Joan found 70 seashells on the beach...") that frontier AI found challenging 12 years prior.
- [00:23] – Graph of "AI Progress is Getting Faster" (Source: AddSub / MATH / Epoch AI) plotting performance trajectories across word number problems, competitive maths, and frontier maths from 2013 to 2026.
- [01:21] – Stanford AI Index chart ("Something Has Changed in AI") comparing slower pre-LLM benchmark progress against rapid vertical climbs in benchmarks such as image classification, reading comprehension, PhD-level science questions, and autonomous software engineering relative to human performance (100%).
- [02:06] – A Reddit screenshot (r/ChatGPT) detailing ChatGPT's unintended obsession with the word "goblin" across conversations following a model update.
- [02:32] – Bar chart titled "Where the Goblins Came From" (Source: OpenAI) showing message frequency changes containing "goblin" between GPT-5.2 and GPT-5.4 across different personality modes (Default, Nerdy, Quirky, etc.).
- [03:11] – Scatter chart titled "Useful Does Not Mean Reliable" (Source: Princeton) plotting Accuracy against Reliability between March 2024 and May 2026.
- [03:43] – Trend line chart titled "At Anthropic, AI No Longer Assists" (Source: Anthropic) tracking the shift of automated tasks from "AI assists" to "AI collaborates" and "AI leads" between August 2025 and August 2026.
- [04:45] – Column chart titled "AI Researchers Are Changing More Code" (Source: OpenAI) showing lines of code changed per active contributor from 2021 through Q3 2026 relative to pre-2025 baselines.
Claims & numbers
- The presenter states that on early word math problems, AI scored approximately 77% in 2013 and reached 94% after several years [00:27].
- On competitive math benchmarks, AI performance rose from 6% to roughly 95% in just a few years [00:43].
- On frontier math benchmarks, AI performance increased from near baseline to 98% in just over a year [01:00].
- An unintended persona/training update between GPT-5.2 and GPT-5.4 caused the rate of messages containing the word "goblin" to spike nearly 4,000% under default settings [02:51].
- Citing a Princeton study, Manthorpe claims that while AI accuracy has steadily scaled up, reliability has remained flat and failed to scale concurrently [03:22].
- Citing Anthropic internal metrics, Manthorpe states that in February 2026, AI led only 1% of tasks on the Claude model, but by August 2026, AI was leading approximately 25% (a quarter) of tasks [04:18].
- Citing OpenAI data, the presenter claims the volume of code changes per active contributor surged to over seven times the pre-2025 baseline by Q3 2026 [04:49].
Notable quotes
- [01:08] – "The problems are getting harder, but the progress is getting faster. This does not feel normal."
- [03:03] – "We don't really totally understand how large language models work, and we don't have a lot of time to test them."
- [05:37] – "I don't think you need to even believe in any of those to think that speed itself is a real problem here."
Assessment
This is a polished television news studio explainer utilizing real empirical charts and industry data releases (OpenAI, Anthropic, Epoch AI, Stanford AI Index, Princeton). The presentation relies on graphic projections and charts rather than live technical benchmarks, framing published industry figures into an overarching narrative about velocity and governance risk.
Described by gemini-3.8-flash on 2026-10-07 from the video's audio and frames.