Hot take on OpenAI GPT-6 Astra, with a challenge to Brockman's AGI claims
Gary Marcus @GaryMarcus · x · 2026-09-03 · ★★★ · archived
The leading LLM skeptic called Astra a genuine advance and a vindication of symbolic world models, while rejecting Greg Brockman's claim that it is AGI.
Summary
Posted on launch day, this thread (with a companion Substack post, garymarcus.substack.com/p/hot-take-on-gpt-6-astra) conceded that Astra "looks to be pretty impressive" and that multiple reports suggest a genuine advance. Marcus said it was vindicating that Astra's ARC-AGI-3 result (63% semi-private, beating humans on 96% of levels) comes from building explicit symbolic models of novel environments, something he has argued for for a decade. He still disputed Brockman's "we're there" AGI framing, predicting problems on open-ended real-world tasks, and flagged that Astra is less monitorable than earlier models. Earlier (Aug 3, x.com/GaryMarcus/status/2084114068248592447) he had argued Astra would be incremental, not a leap. Verified via the X syndication API (2026-09-03 21:34 UTC).
Archived text
Hot take on OpenAI GPT-6 Astra*, with a challenge to @gdb’s claims about it being AGI toward the end:
• Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.
• As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its computations.
• What we don’t know is how robust that capability is. That is THE key question.
• Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.
• And as a scientist, it’s disappointing that we don’t (yet) know much about how the system actually works.
• As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.
• The new system appears to be less monitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability.
• Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK; link: https://garymarcus.substack.com/p/where-will-ai-be-at-the-end-of-2027?r=8tdk6&utm_medium=ios)
——————- *This hot take is VERY tentative, pending more information about how it works and what its limitations are.
Quoting @arcprize: GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:
- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
- It surpasses human performance on 96% of ARC-AGI-3 levels
- It builds the most precise symbolic model of novel environments we've seen
Our analysis:
views 88644 · likes 268 · reposts 23 · replies 22 (at fetch time)
Archived 2026-09-29 via fxtwitter (unofficial).
Related events
- OpenAI releases GPT-6 Astra, its first GPT-6 model 2026-09-03
- GPT-6 Astra scores 62.7% on ARC-AGI-3 (99.9% with provider harness), outacting humans on 96% of levels 2026-09-03
All posts · id: 2026-09-03-marcus-hot-take-gpt-6-astra