DeepSeekMath-V2: open-weights self-verifying prover reaches IMO 2025 gold level and 118/120 on Putnam 2024
DeepSeek released DeepSeekMath-V2 (685B parameters, built on DeepSeek-V3.2-Exp-Base, Apache 2.0). It is trained to write natural-language proofs and check them with an LLM verifier, including a meta-verifier. With scaled test-time compute it reached gold-medal level on IMO 2025 and CMO 2024 and scored 118/120 on Putnam 2024. It was the first openly downloadable model at IMO-gold level.
Key facts
- Paper: arXiv 2511.22570 'DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning' (27 Nov 2025)
- 685B parameters; base DeepSeek-V3.2-Exp-Base; Apache 2.0 weights on Hugging Face
- Gold-level scores on IMO 2025 and CMO 2024; 118/120 on Putnam 2024 (scaled test-time compute)
- Method: faithful LLM proof verifier plus meta-verification to cut hallucinated issues; the generator is rewarded for finding and fixing its own errors; verifier compute is scaled to auto-label hard proofs without human annotation
What happened
Four months after closed models from Google DeepMind and OpenAI reached IMO gold, DeepSeek released open weights for a proof-writing model at the same level. It made self-verification (generator + verifier + meta-verifier) the main training signal instead of final-answer rewards.
Why it matters
It made olympiad-level natural-language proof generation reproducible outside the big US labs. The generate-then-verify recipe became a common pattern in 2026 AI-for-math systems.
Changelog
- 2026-09-29: created
Models
- DeepSeekMath-V2 DeepSeek · current
Related events
- AI systems reach gold-medal level at the International Mathematical Olympiad ★★★★★
- AxiomProver produces machine-checked Lean proofs for all 12 Putnam 2025 problems ★★★
- DeepSeek-V3: frontier-level open model trained for ~$5.6M in GPU time ★★★★★
Sources (2)
id: 2025-11-27-deepseekmath-v2 · updated 2026-09-29 · open in the interactive timeline