MODEL BRIEFING 04 OCT 2026
Meet the new
Kolibri.
Kolibri 1 is a newly released German-English reasoning model from Aleph Alpha. Its Mixture-of-Experts design pairs 78.1 billion total parameters with 3.46 billion active per token, plus tool calling and long-context support.
Open weights · Apache 2.0 terms
REASONING MODEL
THE SHORT VERSION
German by design.
Built to reason.
Kolibri 1 is a German-English Mixture-of-Experts (MoE) reasoning model developed by Aleph Alpha. Released on 3 October 2026, it supports explicit reasoning modes, tool calling and long-context text processing. Its open weights are offered through Hugging Face under Apache 2.0 terms.
Aleph Alpha describes Kolibri as a model for multi-step reasoning, coding, structured extraction, retrieval-augmented generation (RAG), long-document processing and agentic tool workflows. The model card frames its intended use around human-AI collaboration: people should review output before consequential actions are taken.
Specs and intended-use details: Kolibri model card. Launch date and positioning: Aleph Alpha announcement.
A SPARSE MODEL WITH ROOM TO SCALE
Large capacity.
Selective activation.
Kolibri is a 50-layer Transformer Mixture-of-Experts model. Each MoE layer contains 384 experts, while six routed experts and one shared expert are active per token. That keeps per-token computation lower than the full parameter count suggests, while the full weights still require substantial memory.
Read the technical reportDE / EN
TOKEN
¹ Quality and serving efficiency validated up to 1,048,576 tokens. Aleph Alpha recommends up to 262,144 tokens for serving efficiency and complex tasks. Source: model card.
FROM ALEPH ALPHA
European roots.
Open weights.
Aleph Alpha says its teams built Kolibri in Germany and trained it on infrastructure in Germany and Finland, under European and German law. The company presents sovereignty, deployment choice and control over model use as core design goals. Those are the provider’s claims; evaluate the full training, data and deployment details for your own requirements.
German and English
The model targets both languages, with a tokenizer designed for German word structure.
Up to one million tokens
Validated up to 1M tokens; the provider recommends 256K or less for serving efficiency and complex work.
Apache 2.0 terms
Weights are available from Hugging Face. Check the official license and model card before use.
PUBLIC RELEASE
Available to
download now.
Kolibri is publicly available as open weights under Apache 2.0 terms. Running it still takes substantial GPU memory and the vendor-supported inference stack. The one-million-token figure is a validated upper context limit; the model card recommends shorter contexts for efficient serving.
Open the model cardA PRACTICAL DISTINCTION
Low active compute.
High memory needs.
Mixture-of-Experts activation reduces how many parameters are used for each token. It does not remove the need to hold the full model in memory. The BF16 model card lists about 156 GB of model memory and gives multiple-GPU and accelerator configurations; the FP8 release uses less memory. Actual requirements depend on format, context length, batching and serving setup.
Explore Kolibri if you:
need German-English reasoning, document processing, retrieval over your own data, code assistance or tool calling with a human checking results.
Plan carefully if you:
have limited accelerator memory, expect one million tokens to be the efficient default, or need a model to take consequential actions without human review.
Hardware, context recommendations and intended-use notes are from the official model card.
KOLIBRI FAQ
Clear answers,
with sources.
Specifications can change. Follow Aleph Alpha’s official release and model card for updates.
What is Kolibri?+
Kolibri 1 is Aleph Alpha’s German-English Mixture-of-Experts reasoning model, released on 3 October 2026. It supports reasoning modes, tool calling and long-context text tasks.
Can I download Kolibri?+
Yes. Aleph Alpha publishes the full weights on Hugging Face under Apache 2.0 terms. Consult the model card for the available weight variants and usage instructions.
How many parameters are active?+
The model card lists 78.1 billion total parameters and 3.46 billion active parameters per token. The full model still needs to fit in accelerator memory.
Does Kolibri really support a million-token context?+
Aleph Alpha reports validated quality and serving efficiency up to 1,048,576 tokens. For serving efficiency and complex tasks, its model card recommends a context of at most 262,144 tokens.
What can Kolibri do?+
It is designed for German and English multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and tool calling. The provider describes human-reviewed workflows as the intended setting.
What hardware does Kolibri need?+
The BF16 model card reports about 156 GB of model memory. Its listed minimum configurations include four A100 80 GB GPUs, four H100 SXM5 GPUs, two H200 GPUs, one B200 or one B300. FP8 weights use less memory; check the model card for current options and inference requirements.
READ THE ORIGINALS
Good questions
start with good sources.
This independent overview summarizes provider-published information. Product positioning and benchmark results are Aleph Alpha’s own claims; they are not independent evaluations.
- 01Aleph Alpha
Kolibri Has Landed: A Sovereign Open-Weight Model · 3 October 2026
Read launch announcement ↗ - 02Hugging Face model card
Model overview, intended use, context window and hardware requirements
Read the model card ↗ - 03Aleph Alpha
Official Kolibri product and deployment information
Visit Kolibri page ↗ - 04Technical report
Architecture and training details from the model developer
Read technical report ↗