Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. ARC Prize 2026 Milestone #2: open-source ARC-AGI-3 agents…

ARC Prize 2026 Milestone #2: open-source ARC-AGI-3 agents top out at 27.9% under Kaggle compute limits

★★after cutoffbenchmarkARC Prize FoundationTufa Labsconfidence: high

On Oct 1, 2026 ARC Prize announced the Milestone #2 winners of its ARC-AGI-3 Kaggle competition, all of whom open-sourced their agents: Daniel Franzen (27.9%, $25K), "Lord Han Solo" (23.8%, $7.5K) and Lohit Siriki (22.5%, $5K). The open, compute-limited track remains far behind frontier models with provider harnesses (GPT-6 Astra 99.9%, GPT-6.1 Sol 96.4%). Tufa Labs leads the live Kaggle leaderboard with 45.33% (Sept 29), and several winning entries build on its open "Duck" harness.

Key facts

What happened

ARC Prize's Kaggle competition for ARC-AGI-3, its interactive-game reasoning benchmark, gives interim prizes to open-sourced solutions. On Oct 1, 2026 ARC Prize named the Milestone #2 winners. Daniel Franzen scored 27.9% (a Daniel Franzen was also part of "The ARChitects", the ARC Prize 2024 winners; ARC Prize did not say whether this is the same person), Kaggle user "Lord Han Solo" 23.8% and Lohit Siriki 22.5%. All three published their notebooks on Kaggle. Two of the notebooks say they build on Tufa Labs' "Duck" harness, which won Milestone #1. Tufa Labs itself leads the live leaderboard at 45.33% (Sept 29), but it is not among the Milestone #2 prize winners.

Why it matters

The Kaggle track runs open models under Kaggle's compute limits, so it measures how far open methods get without frontier APIs. In the same week, ARC Prize verified frontier API models with provider harnesses near the ceiling: GPT-6 Astra at 99.9% and GPT-6.1 Sol at 96.4%. Open agents under Kaggle's limits are at roughly 22–45%. That gap is the main open question before the final results on Dec 4.

Changelog

  • 2026-10-04: created (outcome of upcoming item 2026-09-30-arc-prize-2026-milestone-2)

Related posts (2)

Related events

  1. ARC Prize launches ARC-AGI-3, an interactive game benchmark where frontier AI scored under 1% ★★★★
  2. GPT-6 Astra scores 62.7% on ARC-AGI-3 (99.9% with provider harness), outacting humans on 96% of levels ★★★★★
  3. OpenAI releases GPT-6.1 Sol: near-Astra performance at one-fifth of Astra's price ★★★★

Sources (8)

id: 2026-10-01-arc-prize-2026-milestone-2-winners · updated 2026-10-04 · open in the interactive timeline