Peter Gostev: "Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them"
Peter Gostev @petergostev · x · 2026-10-06 · ★★ · archived
Widely shared compute comparison (~162k views by 7 Oct); both numbers are rounded from earlier claims, not new disclosures.
Summary
Posted 13:58 UTC on 6 Oct by Peter Gostev (AI capability at Arena, per his X bio), quoting NVIDIA AI Infrastructure ("Nearly 4,000 NVIDIA Grace Blackwell Superchips powered the training of Mistral Large 4"): "Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them. We'll see what the kitten can pull off, but the disparity is pretty stark". Check: Mistral's own figure is 3,800 Grace Blackwell GPUs (blog; Lample). The post gives no source for the Astra figure. It matches earlier claims in this dataset (GPT-6 Astra entry: "more than 100,000 GPUs at the Stargate site in Texas"; Jensen Huang on 6 Sept: "~100K+ NVIDIA Grace Blackwell NVLink72"), not a new OpenAI disclosure. Reach at 11:31 UTC on 7 Oct: ~162k views, 2,903 likes.
Archived text
Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them.
We'll see what the kitten can pull off, but the disparity is pretty stark
Quoting @NVIDIAAIInfra: Frontier scale, built in Europe.
Nearly 4,000 NVIDIA Grace Blackwell Superchips powered the training of Mistral Large 4, now in public preview.
Congrats to the @MistralAI team 🙌
views 161576 · likes 2903 · reposts 108 · replies 56 · quotes 14 (at fetch time)
Archived 2026-10-07 via manual (api.fxtwitter.com, fetched 11:31 UTC).
People
Related events
All posts · id: 2026-10-06-petergostev-mistral-4000-gpus