Post-Cutoff.com
  1. Home
  2. Posts
  3. Hot take on OpenAI GPT-6 Astra, with a challenge to…

Hot take on OpenAI GPT-6 Astra, with a challenge to Brockman's AGI claims

Gary Marcus @GaryMarcus · x · 2026-09-03 · ★★★ · archived

Open the original ↗

The leading LLM skeptic called Astra a genuine advance and a vindication of symbolic world models, while rejecting Greg Brockman's claim that it is AGI.

Summary

Posted on launch day, this thread (with a companion Substack post, garymarcus.substack.com/p/hot-take-on-gpt-6-astra) conceded that Astra "looks to be pretty impressive" and that multiple reports suggest a genuine advance. Marcus said it was vindicating that Astra's ARC-AGI-3 result (63% semi-private, beating humans on 96% of levels) comes from building explicit symbolic models of novel environments, something he has argued for for a decade. He still disputed Brockman's "we're there" AGI framing, predicting problems on open-ended real-world tasks, and flagged that Astra is less monitorable than earlier models. Earlier (Aug 3, x.com/GaryMarcus/status/2084114068248592447) he had argued Astra would be incremental, not a leap. Verified via the X syndication API (2026-09-03 21:34 UTC).

Archived text

Hot take on OpenAI GPT-6 Astra*, with a challenge to @gdb’s claims about it being AGI toward the end:

• Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.

• As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its computations.

• What we don’t know is how robust that capability is. That is THE key question.

• Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.

• And as a scientist, it’s disappointing that we don’t (yet) know much about how the system actually works.

• As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.

• The new system appears to be less monitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability.

• Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK; link: https://garymarcus.substack.com/p/where-will-ai-be-at-the-end-of-2027?r=8tdk6&utm_medium=ios)

——————- *This hot take is VERY tentative, pending more information about how it works and what its limitations are.

Quoting @arcprize: GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:

  • Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
  • It surpasses human performance on 96% of ARC-AGI-3 levels
  • It builds the most precise symbolic model of novel environments we've seen

Our analysis:

views 88644 · likes 268 · reposts 23 · replies 22 (at fetch time)

Archived 2026-09-29 via fxtwitter (unofficial).

Related events

All posts · id: 2026-09-03-marcus-hot-take-gpt-6-astra