innFactory AI Consulting from Rosenheim, Germany adds Kolibri-1 because it is the first language model Aleph Alpha has released under Apache 2.0 (Pharia-1 and T-Free are under the non-commercial Open Aleph License), and one built specifically for German within the class of sparse Mixture-of-Experts models. This page is based on the model card on Hugging Face, the tech report, Aleph Alpha’s announcement and media coverage of the launch. This overview reflects the state as of October 2026.
What is Kolibri-1?
On 3 October 2026, Aleph Alpha Research released Kolibri-1: a German-English Mixture-of-Experts transformer with 78.1 billion parameters, of which 3.46 billion (4.4 percent) are active per token. According to the tech report, the model is built for organisations that process sensitive data under regulation and therefore have to run their models on infrastructure they control: public administration, industry, aerospace. The weights are available under Apache 2.0 on Hugging Face.
| Kolibri-1 | |
|---|---|
| Architecture | MoE, 50 layers, 384 experts + 1 shared expert, 6 active |
| Parameters | 78.1B total, 3.46B active |
| Context | 1,048,576 tokens (native 262,144) |
| Attention | sliding window of 512 tokens, full attention in every 5th layer; 48 query / 4 KV heads |
| Tokenizer | UniBPE, 128,000 entries |
| Languages | German, English |
| Reasoning | none, low, medium, high |
| Weights | FP8 (Kolibri-1), BF16 (Kolibri-1-BF16) |
| Knowledge cutoff | 18 June 2026 |
| Licence | Apache 2.0 |
Kolibri-1 is a new development, not a successor within the Pharia or Luminous stack: the model card mentions neither Pharia nor the T-Free architecture, and the tech report describes a model trained from scratch. Its predecessor was an internal model called Kolibri Origin (30.6 billion parameters, 3.27 billion active) with which Aleph Alpha validated the pipeline; it was not released.
Training: 24 trillion tokens, more than 20 percent German
According to the tech report, training covers about 24 trillion tokens: 20 trillion of pre-training, 3.44 trillion of mid-training and about 200 billion for the context extension to 256k tokens. The German share is 21.3 percent of pre-training; to reach it, Aleph Alpha curated more than 2 trillion German tokens from the web with its own pipeline and generated them synthetically. English makes up about 62 percent, code about 14 percent.
Pre-training ran on 768 NVIDIA B200 GPUs in 96 nodes over 21 days (about 392,000 GPU hours) on infrastructure in Germany and Finland; the model card estimates energy consumption at about 950 MWh including data centre overhead. Post-training combines supervised fine-tuning on curated German and English data with reinforcement learning, among others in retrieval and code environments.
The UniBPE tokenizer with 128,000 entries is designed bilingually. In its blog, Aleph Alpha reports about 4.9 bytes per token on German web text (FineWeb-2) and 4.58 bytes per token on English (FineWeb). For German compounds, that means less fragmentation and therefore fewer tokens per document than with English-centric tokenizers; the actual saving depends on the corpus.
Benchmarks: Aleph Alpha’s figures
In the tech report (Tables 28 and 29), Aleph Alpha compares Kolibri-1 with open MoE models of the same class, i.e. roughly 3 to 6 billion active parameters. These are vendor figures; all models were measured by Aleph Alpha in the same environment.
| Benchmark | Kolibri-1 | Qwen3.5 35B-A3B | Qwen3.6 35B-A3B | Gemma 4 26B-A4B | Nemotron 3 Nano 30B-A3B |
|---|---|---|---|---|---|
| Overall (EN) | 75.5 | 74.7 | 71.4 | 71.9 | 65.6 |
| Overall (DE) | 70.8 | 69.8 | 67.3 | 66.3 | 59.3 |
| GPQA Diamond (DE) | 81.3 | 84.2 | 80.6 | 80.1 | 49.6 |
| MMLU-ProX CoT (DE) | 75.5 | 81.7 | 81.9 | 81.1 | 61.0 |
| AIME 2025 (DE) | 87.5 | 76.7 | 82.9 | 88.1 | 84.4 |
| LiveCodeBench v6 | 85.9 | 77.8 | 82.5 | 82.3 | 71.3 |
| SWE-Bench Verified | 66.4 | 71.6 | 73.8 | 57.8 | 38.6 |
| Tau2-Bench Telecom | 94.7 | 97.7 | 99.1 | 45.3 | 45.9 |
| BFCL v4 (overall) | 61.4 | 70.5 | 67.2 | 68.2 | 61.5 |
| BFCL v3 (multi-turn) | 39.8 | 54.0 | 53.5 | 53.4 | 47.9 |
| RGB Closed-Book | 51.0 | 81.0 | 79.0 | 79.0 | 80.0 |
| AA-Omniscience non-hallucination | 44.0 | 11.1 | 56.7 | 14.3 | 19.0 |
| Industry RAG (DE) | 67.5 | 70.0 | 65.8 | 49.4 | 46.1 |
Three observations. First, on the overall scores Kolibri-1 sits just ahead of the comparison models in its class in both languages, with the largest margin in German. Second, mathematics, code and retrieval-grounded tasks are its strengths: AIME 2025 in English 96.9, Honeypot 80.8, Industry RAG in English 89.7. Third, the results for closed-book factual knowledge (RGB Closed-Book 51.0) and multi-turn tool calling (BFCL v3 multi-turn 39.8, TerminalBench 2.1 27.7) show clear gaps; Aleph Alpha itself positions the model for use with retrieval and describes it as abstaining rather than guessing when the context does not support an answer. The high non-hallucination rate of 44.0 on AA-Omniscience fits this design: the model answers incorrectly less often because it more often does not answer at all.
For long context, the model card reports RULER scores for the base model of 63.2 at 1 million tokens (Qwen3.5 35B: 57.5) but 69.8 at 256k versus 80.1 for Qwen; for efficiency, the model card recommends inputs of up to 262,144 tokens.
Operations: hardware and vLLM
Kolibri-1 is designed for self-hosting. The FP8 weights take about 78 GB; embeddings, LM head, norms and MoE router are in BF16, the KV cache in FP8.
| GPUs | |
|---|---|
| Minimum (model card) | 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300 |
| Recommended (model card) | 2× H100 SXM5, 2× H200, 1× B200 or 1× B300 |
| Concurrent 256k requests in FP8 (tech report, Table 31) | 2× H100: ~18, 4× H100: ~67, 1× B200: ~31, 2× B300: ~157 |
Serving starts via vLLM with Aleph Alpha’s plugin:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choiceAlternatively there is a container image at ghcr.io/aleph-alpha/aleph-alpha-inference. The model card recommends temperature 1.0, top-p 0.97 and top-k 128. The server is OpenAI-compatible, so Kolibri-1 can be run like other self-hosted models as a backend for CompanyGPT or behind the AI Gateway. A hosted Aleph Alpha API for Kolibri-1 is not currently documented; the model card points to sales for deployment and specialisation. WirtschaftsWoche and other media report planned offerings on STACKIT together with the Schwarz Group.
Data protection, regulation and sovereignty
In the tech report, Aleph Alpha describes a development pipeline that takes the EU AI Act, the General-Purpose AI Code of Practice and the GDPR into account from the start: the data pipeline filters against illegal, harmful and pirated content and redacts personal data from sources; the model is meant to abstain when the context does not support an answer. The model card also lists limitations (systemic biases from the training data, knowledge outdated after June 2026, hallucinations) and recommends prompt design, application-level validation, output filters and red-teaming; the model is not recommended for unsupervised high-stakes decisions.
For the sovereignty question, the deployment path matters most: with open weights on your own hardware, neither prompts nor documents leave your environment. That distinguishes Kolibri-1 from hosted models such as OpenAI GPT or Anthropic Claude, where EU data residency is established through cloud deployments such as Microsoft Foundry or Bedrock. Training itself took place in Germany and Finland according to Aleph Alpha.
Positioning: Kolibri-1, Pharia, Mistral and OpenEuroLLM
| Kolibri-1 | Pharia-1 LLM-7B | Mistral (open models) | OpenEuroLLM | |
|---|---|---|---|---|
| Vendor | Aleph Alpha (Heidelberg) | Aleph Alpha | Mistral AI (Paris) | EU consortium |
| Architecture | MoE 78.1B / 3.46B active | dense 7B | dense and MoE, depending on model | in development |
| Context | 1M tokens | 8k tokens | depending on model | – |
| Languages | German, English | multilingual (7 EU languages) | multilingual | all EU languages (goal) |
| Licence | Apache 2.0 | Open Aleph License (non-commercial) | Apache 2.0 (open models) | open (goal) |
| Hosted API | not documented | PhariaAI | Mistral La Plateforme, Azure, AWS and others | – |
Kolibri-1 is the model for organisations that need German as their main language, a long context and operation on their own hardware, and can do without image input and additional languages for that. Anyone looking for broad language coverage or a hosted API is better served by Mistral; anyone waiting for a consortium-backed EU solution keeps an eye on OpenEuroLLM.
Our recommendation
For public authorities, industrial companies and regulated sectors in the DACH region, Kolibri-1 is a candidate for the shortlist of self-hosted models: open licence, German language quality within its class, 1 million tokens of context, reasoning levels and a deployment path via vLLM that gets by with two H100s or one B200. Check three points before deploying: quality in multi-turn tool calling (where the model trails its comparison models), the handling of factual questions without retrieval (the model is designed for RAG) and the operating cost of the GPU infrastructure compared with an EU cloud deployment of a hosted model. For the existing Aleph Alpha portfolio: Kolibri-1 replaces Pharia-1 LLM-7B as our recommendation.
We evaluate Kolibri-1 on a per-project basis as a backend for CompanyGPT and as a model behind the AI Gateway. For an assessment of whether Kolibri-1 fits your requirements, contact innFactory AI Consulting.
