Kolibri Release Guide: Aleph Alpha's 78B Open Model

TL;DR
Kolibri is an Apache 2.0 English-German MoE with 78B total and 3B active parameters and a 1M context. Benchmarks, the vLLM command, and who should run it.
Aleph Alpha released Kolibri on October 3, 2026: an English-German Mixture-of-Experts model with 78B total and about 3B active parameters, a context window of up to 1M tokens, and open weights under the Apache 2.0 license. It is aimed at sovereign, on-premise work in public administration, industry and aerospace, so the interesting questions are not whether it tops a leaderboard (it does not) but where it fits, what it costs to run, and whether its German and grounding claims hold. This guide covers what shipped, the vendor's benchmark numbers with the caveats, the exact serving command, and who should try it. Everything below is from Aleph Alpha's release post unless stated otherwise.
Last updated: October 3, 2026
What shipped#
| Item | Detail |
|---|---|
| Architecture | Mixture-of-Experts Transformer, 78.1B total and 3.46B active parameters per token, 384 experts with 6 active |
| Context | Up to 1M tokens (trained to 256k, extended with a serve-time flag) |
| License | Apache 2.0, full weights on Hugging Face |
| Reasoning | Four effort levels: none, low, medium, high |
| Languages | English and German, with a bilingual tokenizer; 21.3% of pre-training tokens are German |
| Training | 20T pre-training tokens on 768 B200 GPUs over 21 days, then mid-training and long-context stages |
| Hosted API price | None published in the announcement. Enterprise deployment goes through Aleph Alpha sales |
Aleph Alpha says the 78B size was chosen for serving cost: in its tests a 123B variant handled only 3 concurrent 256k-token queries on two H100s, while the 78B handled 18 and decoded 28% faster. Those are the vendor's measurements.
Benchmarks, with the honest read#
These are vendor-published numbers from Aleph Alpha's own evaluation harness, run at each model's highest reasoning effort. Higher is better.
| Benchmark | Kolibri | Qwen3.6-35B-A3B | Nemotron 3 Super 120B-A12B |
|---|---|---|---|
| AIME 2026 | 96.0 | 91.0 | 90.4 |
| GPQA Diamond | 84.3 | 83.4 | 78.0 |
| LiveCodeBench v6 | 85.9 | 82.5 | 82.0 |
| SWE-Bench Verified | 66.4 | 73.8 | 60.2 |
| BFCL v4 overall | 61.4 | 67.2 | 61.0 |
| Tau2-Bench Telecom | 94.7 | 99.1 | 68.1 |
What the table says: Kolibri leads on math and competitive coding, and is mixed on tool use and software-engineering tasks, where Qwen3.6 is ahead. The release post's larger table also lists Qwen3.8 27B (a dense model) ahead overall, 80.2 against Kolibri's 75.5 in English, so this is not the strongest open model per benchmark. Aleph Alpha's own framing is quality per serving cost, claiming Kolibri sits on the Pareto frontier among the models it compared, which is a claim about efficiency, not about the top score.
On grounding, the vendor's headline is that Kolibri is trained to say "I don't know" when the context does not contain the answer. On the public AA-Omniscience set its non-hallucination rate is 44.0%, up from 14.8% for its predecessor, but Qwen3.6-35B-A3B scores 56.7% in the same table. Treat the abstention story as a real improvement over the predecessor, not a lead over every peer.
Run it (from the official post)#
These commands come straight from Aleph Alpha's instructions. We have not run them.
pip install "aleph-alpha-inference>=1.0"
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
A container image is also provided: ghcr.io/aleph-alpha/aleph-alpha-inference. To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'. The recommended sampling is temperature=1.0, top_p=0.97 and top_k=128. Kolibri needs Aleph Alpha's vLLM plugin, shipped in the aleph-alpha-inference package, so plan on that package rather than a stock vLLM install.
What people are saying#
- Speed is real, thinking is long. One commenter ran it on a single RTX Pro 6000 in fp8 and measured about 170 tokens per second, which they attribute to the 3B active parameters, but found it spends a lot of tokens overthinking even when it reaches the right approach (Hacker News).
- German documents. Another commenter's read is that it is good at reading German documents and reasoning over them, with a custom German tokenizer and low hallucination (Hacker News).
- Open weights versus an open pipeline. Several commenters argued that sovereignty needs more than weights: open data and a reproducible training pipeline, not just a download (Hacker News).
- Is sovereignty the right frame? One commenter asked why sovereign models matter if open-source models exist, suggesting hosting is the real dependency, and a German commenter said they were more excited by Mistral and Black Forest Labs (Hacker News).
Who should try it#
- German-language document work on your own hardware. Public sector, legal or industrial text where data cannot leave your infrastructure is exactly what the model targets, and the bilingual tokenizer means fewer tokens per German task.
- Agent builders who want a fast, cheap-to-serve open model. A 3B-active model is quick, but test tool calling yourself: the vendor's own numbers are behind Qwen3.6 on BFCL and SWE-Bench.
- Everyone else: if you want the strongest open coding model, compare it first against the options in our open-weights coding showdown and best local coding LLMs. For the other European open model effort, see our Apertus explainer.
FAQ#
What is Kolibri?#
Aleph Alpha's open-weight English-German Mixture-of-Experts language model, released October 3, 2026 under Apache 2.0, with 78B total and about 3B active parameters.
Can I use Kolibri commercially?#
The weights are published under Apache 2.0, which permits commercial use, per the release post. Check the license file on Hugging Face for the exact terms before you ship.
What hardware do I need?#
Aleph Alpha does not publish a minimum in the announcement. One Hacker News commenter reports running it in fp8 on a single RTX Pro 6000. Plan to test on your own hardware.
Is Kolibri good at coding?#
Mixed. It scores 85.9 on LiveCodeBench v6 but 66.4 on SWE-Bench Verified, behind Qwen3.6-35B-A3B at 73.8 in the vendor's table.
Continue Reading#
- Apertus: Europe's Sovereign Open Model - the other European open-weight effort
- Best Local Coding LLMs 2026 - what to run on your own machine
- Local LLM Runtimes for Coding Agents - serving options once you pick a model
- GLM 5.2 vs DeepSeek V4 vs Qwen3 - the open coding models Kolibri is measured against
- Kimi K3 Open Weights on Hugging Face - another recent open release and its serving story
Sources#
- Kolibri Has Landed: A Sovereign Open-Weight Model: Aleph Alpha's release post, the source for every spec, benchmark and command (primary)
- Aleph-Alpha/Kolibri-1: the weights, linked from the release post (not independently read)
- Hacker News: Show HN: Germany's new sovereign AI model Kolibri: community discussion (summarized from a read of the thread)
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on local and open-weight models
Apertus: Europe's Answer to AI Sovereignty - and Why HN Is Skeptical
Switzerland's fully open foundation model promises transparent training data and EU compliance. The HN crowd has questions about actual performance.
6 min readThe Best Local Coding LLMs in 2026: Run Enterprise-Grade AI Without the Cloud
Choosing a local coding LLM in 2026 means balancing benchmark performance, hardware cost, and the compliance pressure to keep code off third-party servers. Here is what to run and on what hardware.
8 min readOllama vs LM Studio vs vLLM vs llama.cpp: Picking a Local Runtime for Coding Agents
A fair, sourced comparison of the four runtimes developers reach for when they want a coding agent talking to a model on their own hardware instead of an API: Ollama's convenience, LM Studio's GUI, vLLM's throughput, and llama.cpp's control. What each is actually for, and which to pick.
10 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








