GPT-6 Astra agent wins NetHack: first recorded LLM-agent ascension
On Sept 21, 2026 a GPT-6 Astra agent ascended (won) NetHack 3.6.7 on the public Hardfought server after 37,140 turns, on its third attempt. Hobbyist developer "Kenny" (kenforthewin) ran it and says it is the first recorded ascension by an LLM agent. Astra wrote its own terminal harness, and the run was open-book with human supervision, so it is not a benchmark score. NetHack had been a standard hard long-horizon RL/LLM test since the NetHack Learning Environment (2020) and BALROG (2024).
Key facts
- Ascension: Sept 21, 2026, NetHack 3.6.7 on Hardfought, dwarven Valkyrie 'CodexDelver', 37,140 game turns; third run, played over about twelve days (Sept 9–21) with pauses
- Model: GPT-6 Astra (the author's subscription plus ~$300 of extra credits; no fixed compute budget tracked)
- Astra built its own harness: a 144×36 annotated terminal view, guarded movement batching, a route planner, a Sokoban solver, inventory helpers, and persistent Markdown/JSON notes. The author: 'I didn't write a single line of code, and honestly I didn't even read Astra's code'
- Caveats from the author: open book (NetHack wiki, spoilers, game source code, web); a human set up infrastructure, made suggestions, and stopped and resumed sessions; tools and instructions changed during play; 'This run is not a BALROG submission'
- For comparison, BALROG's leaderboard (Sept 18, 2026) showed GPT-6-Astra-Max at 13.2 ± 2.7% average NetHack progress under its standard protocol. The author's best January 2026 run (other models, handmade harness) reached dungeon level 10
- Evidence: Hardfought server dumplog; harness released as MIT-licensed code with an evidence archive (checksums, redaction manifest)
- Reaction: NLE co-creator Tim Rocktäschel (UCL, ex-DeepMind) on Sept 25: 'Aaaaaaaand it's apparently solved', quoting his post from days earlier that NetHack 'still seems far from being solved' (~90k views)
- A pull request to the NetHackers community site (dunnolab, PR #114) adds it as 'the first LLM-agent ascension' while warning against overstating it
What happened
NetHack is a 1987 roguelike with permadeath and huge item and monster interactions, and a full game takes tens of thousands of turns. It became a standard test of long-horizon agents through Meta/UCL's NetHack Learning Environment (2020) and the BALROG LLM-agent benchmark (2024). Before this run, no AI agent was known to have won it.
A hobbyist who blogs as "Kenny" (kenforthewin) had tried since January to build a harness for LLMs to play NetHack. His best early runs reached dungeon level 10. With GPT-6 Astra he let the model build its own harness. Its first run came close, and its third run ascended on Sept 21, 2026 after 37,140 turns. He published the dumplog, the code and a reviewed evidence archive.
Why it matters
The run shows long-horizon execution, not only game knowledge: weeks of play in which one mistake can end the game. The model also built its own tools and memory along the way. The result is anecdotal. It was open book, supervised by a human, without a fixed budget, and a single success, while Astra's standardized BALROG score was still about 13%. A protocol-controlled ascension on a benchmark would be the next step.
Changelog
- 2026-10-02: created (from leads queue; Zvi AI #188 / Tim Rocktäschel post)
Related posts (1)
- Tim Rocktäschel original ↗ Tim Rocktäschel @_rockt · x · 2026-09-25
Cited as a source by: 2026-09-21-gpt-6-astra-nethack-ascension
Related events
- OpenAI releases GPT-6 Astra, its first GPT-6 model ★★★★★
- GPT-6 Astra scores 62.7% on ARC-AGI-3 (99.9% with provider harness), outacting humans on 96% of levels ★★★★★
- GPT-6 Astra breaks an unsolved 1809 Napoleonic cipher letter to Marshal Marmont from a single scan ★★★
Sources (7)
- officialKenny (Vaguely Aligned blog): An LLM Beat NetHack
- officialHardfought dumplog of the ascension (CodexDelver)
- codeGitHub: kenforthewin/nethack_astra (agent-built harness)
- discussiondunnolab/nethackers PR #114: record the first LLM-agent ascension
- discussionTim Rocktäschel on X: 'Aaaaaaaand it's apparently solved'
- pressKorben: An AI barely managed to finish NetHack
- discussionKenny: It's 2026. Can LLMs play NetHack yet? (January baseline)
id: 2026-09-21-gpt-6-astra-nethack-ascension · updated 2026-10-02 · open in the interactive timeline