Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. GPT-6 Astra agent wins NetHack: first recorded LLM-agent…

GPT-6 Astra agent wins NetHack: first recorded LLM-agent ascension

★★★after cutoffmilestoneOpenAIconfidence: high

On Sept 21, 2026 a GPT-6 Astra agent ascended (won) NetHack 3.6.7 on the public Hardfought server after 37,140 turns, on its third attempt. Hobbyist developer "Kenny" (kenforthewin) ran it and says it is the first recorded ascension by an LLM agent. Astra wrote its own terminal harness, and the run was open-book with human supervision, so it is not a benchmark score. NetHack had been a standard hard long-horizon RL/LLM test since the NetHack Learning Environment (2020) and BALROG (2024).

Key facts

What happened

NetHack is a 1987 roguelike with permadeath and huge item and monster interactions, and a full game takes tens of thousands of turns. It became a standard test of long-horizon agents through Meta/UCL's NetHack Learning Environment (2020) and the BALROG LLM-agent benchmark (2024). Before this run, no AI agent was known to have won it.

A hobbyist who blogs as "Kenny" (kenforthewin) had tried since January to build a harness for LLMs to play NetHack. His best early runs reached dungeon level 10. With GPT-6 Astra he let the model build its own harness. Its first run came close, and its third run ascended on Sept 21, 2026 after 37,140 turns. He published the dumplog, the code and a reviewed evidence archive.

Why it matters

The run shows long-horizon execution, not only game knowledge: weeks of play in which one mistake can end the game. The model also built its own tools and memory along the way. The result is anecdotal. It was open book, supervised by a human, without a fixed budget, and a single success, while Astra's standardized BALROG score was still about 13%. A protocol-controlled ascension on a benchmark would be the next step.

Changelog

  • 2026-10-02: created (from leads queue; Zvi AI #188 / Tim Rocktäschel post)

Related posts (1)

Related events

  1. OpenAI releases GPT-6 Astra, its first GPT-6 model ★★★★★
  2. GPT-6 Astra scores 62.7% on ARC-AGI-3 (99.9% with provider harness), outacting humans on 96% of levels ★★★★★
  3. GPT-6 Astra breaks an unsolved 1809 Napoleonic cipher letter to Marshal Marmont from a single scan ★★★

Sources (7)

id: 2026-09-21-gpt-6-astra-nethack-ascension · updated 2026-10-02 · open in the interactive timeline