Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. GPT-6 Astra scores 62.7% on ARC-AGI-3 (99.9% with provider…

GPT-6 Astra scores 62.7% on ARC-AGI-3 (99.9% with provider harness), outacting humans on 96% of levels

★★★★★after cutoffbenchmarkARC Prize FoundationOpenAIconfidence: high

ARC Prize reported on 2026-09-03 that OpenAI's GPT-6 Astra scored 62.7% on ARC-AGI-3 (semi-private) with the standard harness ($26K) and 99.9% ($19K) with OpenAI's own provider-adapter harness, using fewer actions than the human baseline on 96% of levels; ARC Prize will now label both conditions separately.

Key facts

What happened

Six months after ARC-AGI-3 launched with frontier models near 0%, GPT-6 Astra reached 62.7% under the neutral harness. With OpenAI's context-management setup it reached 99.9%, a result the shared harness did not reproduce, so ARC Prize now reports both. ARC Prize said Astra "builds the most precise symbolic model of novel environments we've seen."

Why it matters

ARC-AGI-3 was meant to measure human-like skill acquisition; its near-saturation (and the harness gap) shows both how fast agentic reasoning improved in 2026 and how much scaffolding now drives scores.

Changelog

  • 2026-09-29: added post link(s) (OpenAI cluster post research)
  • 2026-09-29: created

Related posts (4)

Related events

  1. ARC Prize launches ARC-AGI-3, an interactive game benchmark where frontier AI scored under 1% ★★★★
  2. Jensen Huang declares "AGI has arrived" with GPT-6 Astra; Greg Brockman: "we're now moving into the AGI era" ★★★★

Sources (5)

id: 2026-09-03-arc-agi-3-gpt-6-astra · updated 2026-09-29 · open in the interactive timeline