Aleph Alpha releases Kolibri, a 78B-parameter (3.5B active) open-weight English–German 'sovereign' MoE model under Apache 2.0
On Oct 3, 2026 (German Unity Day) Aleph Alpha released Kolibri-1, an English–German mixture-of-experts model with 78B total and ~3.46B active parameters, a 1M-token context and a German-optimised tokenizer, as open weights under Apache 2.0. It was trained on ~24T tokens on 768 B200 GPUs in Germany and Finland. It came two weeks after the lab signed its deal to combine with Cohere, and it topped Hacker News (two threads, ~360–400 points each).
Key facts
- Architecture: MoE, 78B total / ~3.46B active parameters, 384 experts per layer with 6 routed per token, 50 layers; sliding-window (512) attention with full attention every 5th layer
- Context: up to 1,048,576 tokens (262,144 recommended; longest trained context 262,144); knowledge cutoff 18 June 2026 (EN/DE)
- Training: ~24T tokens (≈62% English, 14% code, 21.3% / ~4.3T German), curated from >200T raw tokens; 768 Nvidia B200 GPUs in Germany and Finland
- Tokenizer: bilingual EN–DE, 128k vocabulary; Aleph Alpha claims 11.2% fewer German tokens than GPT-5's tokenizer (an independent test on the Basic Law found ~15%)
- Aleph Alpha's benchmarks: AIME 2025 96.9 (EN) / 87.5 (DE); GPQA Diamond 84.3 (EN) / 81.3 (DE); HumanEval+ 92.7; compared against Qwen3.6-35B, Nemotron 3 Super 120B, Mistral Small 4, Gemma and GLM models
- Reasoning mode with configurable effort, tool calling, abstention training (Merlin–Arthur 'M/A grounding score' 0.23)
- Weaknesses reported by reviewers: closed-book knowledge, long multi-turn tool use and agentic coding trail Qwen models
- Weights: huggingface.co/Aleph-Alpha/Kolibri-1 (Apache 2.0); served via Aleph Alpha's aleph-alpha-inference package on vLLM; needs ~2×80GB GPUs
What happened
Aleph Alpha published Kolibri-1 on German Unity Day, Oct 3, 2026, as an Apache-2.0 open-weight model on Hugging Face. It is a sparse MoE (78B total, ~3.5B active) built for German and English enterprise and public-sector use. Aleph Alpha says sovereignty "combines two dimensions: how we built the model, and how it transfers to our customers": the model was trained only on infrastructure in Germany and Finland, and customers can run the weights themselves. A tech report is linked from the blog post.
Aleph Alpha's own benchmark tables put Kolibri ahead of open models of similar active size on maths (AIME 2025 96.9% in English) and GPQA Diamond (84.3%). Independent write-ups (Tejas Kumar) say it is weaker at closed-book recall, long multi-turn tool use and coding agents. All benchmark numbers are Aleph Alpha's own.
Why it matters
Kolibri is a notable European open-weight release, from a lab that agreed in Sept 2026 to combine with Cohere. A model with 3.5B active parameters, 1M context and an Apache licence gives German-speaking governments and companies a capable model they can run themselves.
Changelog
- 2026-10-03: created (sweep: HN front page, two threads)
Models
- Kolibri-1 Aleph Alpha · current
Related events
Sources (6)
- officialAleph Alpha: Kolibri has landed — a sovereign open-weight model
- codeHugging Face: Aleph-Alpha/Kolibri-1 model card
- discussionTejas Kumar: Aleph Alpha Kolibri — how the sovereign German LLM works
- discussionHacker News discussion (official post)
- discussionHacker News discussion (Tejas Kumar explainer)
- pressCrypto Briefing: Aleph Alpha releases Kolibri, a 78B-parameter open-weight AI model built in Europe
id: 2026-10-03-aleph-alpha-kolibri-open-weights · updated 2026-10-03 · open in the interactive timeline