Kolibri-1
Reasoning mode with configurable effort and tool calling. Recommended sampling temperature 1.0, top_p 0.97, top_k 128. Needs ~2×80GB GPUs (min 2× A100 80GB). Benchmarks are Aleph Alpha's own (AIME 2025 96.9 EN, GPQA Diamond 84.3 EN). Weaker than Qwen models on closed-book recall and agentic coding per reviewers. No hosted API price found.
- Context window
- 1,048,576 tokens
- Knowledge cutoff
- 2026-06
- Input
- text
- Output
- text
- License
- apache-2.0
- Verified
- 2026-10-03
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | — | huggingface.co/Aleph-Alpha/Kolibri-1 | — |
| Self-hosted (vLLM + aleph-alpha-inference, OpenAI-compatible API) | — | — | docs |
Notable capabilities (3)
- German-optimised tokenizer: Bilingual EN–DE 128k tokenizer; Aleph Alpha reports 11.2% fewer German tokens than GPT-5's tokenizer (an independent test on the German Basic Law found ~15%). source
- Sparse 1M-token context: 78B-total / ~3.46B-active MoE (384 experts, 6 routed) with 512-token sliding-window attention and a global layer every 5th layer; up to 1,048,576 tokens of context (262,144 recommended). source
- Abstention / grounding training: Trained with a Merlin–Arthur protocol to abstain rather than hallucinate; Aleph Alpha reports an M/A grounding score of 0.23 (scale 0–0.5). source
Aleph Alpha's open-weight English–German MoE. Serve with vLLM plus the aleph-alpha-inference package (reasoning and tool-calling parsers enabled) for an OpenAI-compatible endpoint.
Timeline entry
- Aleph Alpha releases Kolibri, a 78B-parameter (3.5B active) open-weight English–German 'sovereign' MoE model under Apache 2.0 ★★★
On Oct 3, 2026 (German Unity Day) Aleph Alpha released Kolibri-1, an English–German mixture-of-experts model with 78B total and ~3.46B active parameters, a 1M-token context and a German-optimised tokenizer, as open weights under Apache 2.0. It was trained on ~24T tokens on 768 B200 GPUs in Germany…