Post-Cutoff.com
  1. Home
  2. Posts
  3. Peter Gostev: "Mistral Large 4 was trained on 4,000 GPUs…

Peter Gostev: "Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them"

Peter Gostev @petergostev · x · 2026-10-06 · ★★ · archived

Open the original ↗

Widely shared compute comparison (~162k views by 7 Oct); both numbers are rounded from earlier claims, not new disclosures.

Summary

Posted 13:58 UTC on 6 Oct by Peter Gostev (AI capability at Arena, per his X bio), quoting NVIDIA AI Infrastructure ("Nearly 4,000 NVIDIA Grace Blackwell Superchips powered the training of Mistral Large 4"): "Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them. We'll see what the kitten can pull off, but the disparity is pretty stark". Check: Mistral's own figure is 3,800 Grace Blackwell GPUs (blog; Lample). The post gives no source for the Astra figure. It matches earlier claims in this dataset (GPT-6 Astra entry: "more than 100,000 GPUs at the Stargate site in Texas"; Jensen Huang on 6 Sept: "~100K+ NVIDIA Grace Blackwell NVLink72"), not a new OpenAI disclosure. Reach at 11:31 UTC on 7 Oct: ~162k views, 2,903 likes.

Archived text

Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them.

We'll see what the kitten can pull off, but the disparity is pretty stark

Quoting @NVIDIAAIInfra: Frontier scale, built in Europe.

Nearly 4,000 NVIDIA Grace Blackwell Superchips powered the training of Mistral Large 4, now in public preview.

Congrats to the @MistralAI team 🙌

views 161576 · likes 2903 · reposts 108 · replies 56 · quotes 14 (at fetch time)

Archived 2026-10-07 via manual (api.fxtwitter.com, fetched 11:31 UTC).

People

Guillaume Lample Jensen Huang

Related events

All posts · id: 2026-10-06-petergostev-mistral-4000-gpus