Announced 3 Oct 2026 · Sources checked
What did Aleph Alpha release?
Heidelberg-based Aleph Alpha published Kolibri on 3 October 2026, the Day of German Unity, and followed with a press release on 5 October. The model card lists the release date as 3 October. Kolibri is a text-in, text-out language model for German and English that supports reasoning, native tool calling and retrieval-augmented generation.
The company describes Kolibri as an English–German mixture-of-experts transformer with 78.1 billion total parameters, of which about 3.46 billion are active per token. Its technical report puts that at 4.4% of the parameters. If the architecture is unfamiliar, our mixture-of-experts explainer covers why a large sparse model can cost far less to serve than a dense model of the same size.
Kolibri follows Kolibri Origin, an earlier 30B-total, 3B-active model with a 65K-token context window that Aleph Alpha says it used to validate its training pipeline. The new model extends the context to up to 1 million tokens.
How was Kolibri built and trained?
According to the technical report, Kolibri was trained on 24 trillion tokens across pre-training, mid-training and long-context extension. Aleph Alpha says it added more than 2 trillion German tokens, curated from the web with its own pipeline and generated synthetically, so that German makes up more than a fifth of the data. The company’s materials give slightly different figures for that share: the blog cites 21.3% of pre-training tokens, the press release about 23%. Aleph Alpha says translated text accounts for only about 6% of the data.
Post-training combined supervised fine-tuning on curated German and English data with reinforcement learning on more than 1.2 million internal tasks covering reasoning, tool use, instruction following, code and retrieval. The model card lists pre-training on 768 NVIDIA B200 GPUs for 21 days, about 392,000 GPU hours, plus shorter mid-training and long-context phases.
Two design choices stand out. Kolibri uses a bilingual tokenizer tuned for German morphology, intended to represent German text with fewer tokens. Aleph Alpha also says it trained the model with abstention data and a protocol it calls Merlin-Arthur, so that in RAG settings it answers “I don’t know” when the supplied documents do not support an answer.
What do the published benchmarks show?
Aleph Alpha’s blog compares Kolibri with Kolibri Origin, Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B. The table below reproduces a selection of the company’s reported numbers. They are vendor-run evaluations, not independent results.
The pattern is mixed rather than dominant. Kolibri leads on several math and agentic tests but trails Qwen3.6-35B-A3B on τ²-bench telecom, BFCL v4 and long-context LongBench Pro. Its grounding score on the AA-Omniscience Index, a metric where negative values reflect more confident wrong answers, is worse than Qwen’s and Mistral’s in the same table. The company’s central claim is efficiency: it says Kolibri sits on the Pareto frontier of quality versus serving throughput in both English and German among the models it compared.
| Benchmark | Kolibri | Qwen3.6-35B-A3B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|
| AIME 2025 | 96.9 | 84.6 | 91.7 | 79.8 |
| GPQA Diamond (German) | 81.3 | 80.6 | 76.6 | 72.9 |
| τ²-bench telecom | 94.7 | 99.1 | 68.1 | 41.5 |
| LiveCodeBench v6 | 85.9 | 82.5 | 82.0 | 71.2 |
| LongBench Pro | 64.5 | 70.8 | 62.9 | 56.4 |
| AA-Omniscience Index (−100 to 100) | −32.8 | −15.3 | −36.5 | −24.0 |
How can developers run Kolibri?
The model is available as Aleph-Alpha/Kolibri-1 on Hugging Face. At the time of checking it was not gated and was tagged Apache 2.0. Serving requires Aleph Alpha’s aleph-alpha-inference package, which provides a Kolibri plugin for vLLM and exposes an OpenAI-compatible API, or the company’s container image.
The model card lists a memory footprint of about 78 GB for the FP8 weights. Minimum hardware is two A100 80 GB or two H100 GPUs, or a single H200, B200 or B300. That is data-centre hardware, not a laptop. Teams weighing hosted APIs against running it themselves can use our guide to managed versus self-hosted AI and our quantization explainer to estimate what FP8 weights mean in practice.
Aleph Alpha also says it ran internal customer-proxy evaluations for the German public sector, aviation, manufacturing, semiconductors and automotive work, and reports gains across successive training runs. Those suites are not public, so outsiders cannot check them.
Why does a sovereign German model matter?
Aleph Alpha frames Kolibri around “sovereignty”: models developed and trained in Europe, deployable on infrastructure the customer controls, with documented data provenance. For public bodies and regulated companies, the ability to run the weights on premises may matter more than a few benchmark points, because it changes where prompts and documents travel. Our guide to AI data residency explains the questions buyers usually ask.
The press release says all training data was screened against a blocklist of more than 4.5 million URLs, including sources from the European Commission’s Piracy Watch List, and that third-party datasets were checked for license terms and opt-outs. The model card says Aleph Alpha is a signatory of the EU General-Purpose AI Code of Practice. These are the company’s descriptions of its own process.
The release also comes at a moment of corporate change. Aleph Alpha says it has signed an agreement with Cohere to form a transatlantic sovereign AI company, subject to regulatory approval, and that it continues to operate independently until closing.
What are the limitations and open questions?
“Open weights” is not the same as open source. Aleph Alpha has published the weights and configuration under Apache 2.0 and a technical report, but not the full training code or datasets. That distinction is explained in our piece on open weights versus open source.
Kolibri supports only German and English, deliberately. The model card lists risks including systemic and political bias, outdated world knowledge, and other errors, and the abstention behaviour should be tested on a team’s own documents rather than assumed. Finally, the efficiency comparison depends on the hardware and serving setup Aleph Alpha chose; real throughput will vary.
Common questions
Is Kolibri free to use commercially?
The Hugging Face repository is tagged Apache 2.0, a permissive license that allows commercial use. Review the license file and the model card’s terms before deployment.
Does Kolibri run on a single consumer GPU?
Not according to the model card. The FP8 weights need about 78 GB, and the listed minimum is two A100 80 GB or H100 GPUs, or one H200, B200 or B300.
Are the benchmark numbers independent?
No. The scores and the Pareto-frontier claim come from Aleph Alpha’s own evaluations. Independent reproduction had not been published when we checked on 7 October 2026.
What to remember
Kolibri is a credible option for teams that need a permissively licensed German–English model they can host themselves. Treat its benchmark lead as a vendor claim, and test German quality, abstention and serving cost on your own workload before adopting it.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





