↓ Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM Aleph Alpha Germany

Aleph Alpha Kolibri-1

Kolibri-1 by Aleph Alpha: open German-English MoE model (78.1B parameters, 1M tokens, Apache 2.0) for self-hosting with vLLM, hardware needs and GDPR.

License Apache 2.0
GDPR Hosting Available
Context 1048576 Tokens
Modality Text → Text

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
Kolibri-1 (Aleph-Alpha/Kolibri-1, FP8) Recommended
3 October 2026
German-English Mixture-of-Experts transformer: 78.1B parameters, 3.46B active per token, 50 MoE layers with 384 experts plus one shared expert, 6 active experts Context up to 1,048,576 tokens (natively trained to 262,144), hybrid attention of sliding window (512 tokens) and full attention in every fifth layer; knowledge cutoff 18 June 2026 Reasoning at four effort levels (none, low, medium, high), tool calling with structured output, vLLM parsers for reasoning and tool calls in the aleph-alpha-inference package Per the tech report, overall 75.5 (EN) and 70.8 (DE); GPQA Diamond 84.3 (EN) / 81.3 (DE), AIME 2025 96.9 (EN), LiveCodeBench v6 85.9, Tau2-Bench Telecom 94.7, Honeypot 80.8 Trained on about 24 trillion tokens with a German share above 20%, on 768 NVIDIA B200 GPUs in Germany and Finland; UniBPE tokenizer with 128,000 entries designed for German compounds Developed, per Aleph Alpha, with the EU AI Act, the GPAI Code of Practice and the GDPR in mind: filtering of illegal and pirated content, redaction of personal data, abstention when the context does not support an answer Apache 2.0, FP8 weights of about 78 GB, FP8 KV cache; per the tech report about 18 concurrent 256k requests on two H100s
Text only, no image or audio input; German and English only Closed-book knowledge weaker than comparable models: RGB Closed-Book 51.0 (comparison models 78 to 89), AA-Omniscience accuracy 14.8 (tech report); Aleph Alpha recommends deployment with retrieval Multi-turn tool calling below the class average: BFCL v3 multi-turn 39.8 (Qwen3.5 35B-A3B 54.0, Nemotron 3 Nano 58.2); TerminalBench 2.1 27.7 versus 39.7 for Qwen3.5 35B-A3B Needs at least two 80 GB GPUs or one H200/B200; serving requires vLLM with the Kolibri plugin, no offering on Azure, AWS or Google documented at launch All benchmarks come from the vendor; independent evaluations were not available at launch Company is in the middle of the Cohere merger (agreement of 16 September 2026, closing pending); no roadmap for further Kolibri sizes published
Current
Kolibri-1-BF16 (Aleph-Alpha/Kolibri-1-BF16)
3 October 2026
Full BF16 precision for your own fine-tuning, quantization or evaluation Same architecture and licence as Kolibri-1
Roughly twice the memory footprint of the FP8 weights For production, the model card recommends the FP8 variant with FP8 KV cache
Current

Use Cases

Typical applications for this model

Public administration and government agencies
Industry, aerospace, automotive suppliers
Retrieval-augmented generation over internal documents
Agentic workflows with tool calling
Document processing and drafting in German
Internal knowledge and research tools

Technical Details

API, features and capabilities

API & Availability
Availability Open weights (Hugging Face); self-hosting with vLLM via the aleph-alpha-inference package (OpenAI-compatible server), container ghcr.io/aleph-alpha/aleph-alpha-inference; no hosted API documented
Latency (TTFT) depends on your own hardware
Features & Capabilities
Tool Use Function Calling Structured Output Reasoning Mode
Training & Knowledge
Knowledge Cutoff 18 June 2026
Fine-Tuning Available (Fine-tuning of the open weights (BF16 variant available), Specialisation and deployment with Aleph Alpha (on request per the model card))
Language Support
Best Quality German, English
Supported 2 languages (German, English)
Trained bilingually with a German share above 20%; tokenizer at about 4.9 bytes per token on German web text (FineWeb-2) per Aleph Alpha

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-Hosted
Own infrastructure or EU cloud
Intended deployment path: Apache 2.0 weights with vLLM (aleph-alpha-inference), about 78 GB GPU memory in FP8; no data leaves your environment
STACKIT (Schwarz Digits)
Germany / Austria
According to media reports at launch, offerings based on Kolibri are planned on STACKIT; no product is currently documented
License & Hosting
License Apache 2.0
Security Filters Model-side (abstention when the context does not support an answer, safety alignment); additional application-level filters recommended
Enterprise Support Yes
On-Premise

Benchmarks

Performance comparison with standardized tests

Overall German (tech report, Aleph Alpha figure)
70.8
Overall English (tech report, Aleph Alpha figure)
75.5
GPQA Diamond German (Aleph Alpha figure)
81.3
AIME 2025 English (Aleph Alpha figure)
96.9
LiveCodeBench v6 (Aleph Alpha figure)
85.9
RULER at 1M tokens, base model (Aleph Alpha figure)
63.2

innFactory AI Consulting from Rosenheim, Germany adds Kolibri-1 because it is the first language model Aleph Alpha has released under Apache 2.0 (Pharia-1 and T-Free are under the non-commercial Open Aleph License), and one built specifically for German within the class of sparse Mixture-of-Experts models. This page is based on the model card on Hugging Face, the tech report, Aleph Alpha’s announcement and media coverage of the launch. This overview reflects the state as of October 2026.

What is Kolibri-1?

On 3 October 2026, Aleph Alpha Research released Kolibri-1: a German-English Mixture-of-Experts transformer with 78.1 billion parameters, of which 3.46 billion (4.4 percent) are active per token. According to the tech report, the model is built for organisations that process sensitive data under regulation and therefore have to run their models on infrastructure they control: public administration, industry, aerospace. The weights are available under Apache 2.0 on Hugging Face.

Kolibri-1
ArchitectureMoE, 50 layers, 384 experts + 1 shared expert, 6 active
Parameters78.1B total, 3.46B active
Context1,048,576 tokens (native 262,144)
Attentionsliding window of 512 tokens, full attention in every 5th layer; 48 query / 4 KV heads
TokenizerUniBPE, 128,000 entries
LanguagesGerman, English
Reasoningnone, low, medium, high
WeightsFP8 (Kolibri-1), BF16 (Kolibri-1-BF16)
Knowledge cutoff18 June 2026
LicenceApache 2.0

Kolibri-1 is a new development, not a successor within the Pharia or Luminous stack: the model card mentions neither Pharia nor the T-Free architecture, and the tech report describes a model trained from scratch. Its predecessor was an internal model called Kolibri Origin (30.6 billion parameters, 3.27 billion active) with which Aleph Alpha validated the pipeline; it was not released.

Training: 24 trillion tokens, more than 20 percent German

According to the tech report, training covers about 24 trillion tokens: 20 trillion of pre-training, 3.44 trillion of mid-training and about 200 billion for the context extension to 256k tokens. The German share is 21.3 percent of pre-training; to reach it, Aleph Alpha curated more than 2 trillion German tokens from the web with its own pipeline and generated them synthetically. English makes up about 62 percent, code about 14 percent.

Pre-training ran on 768 NVIDIA B200 GPUs in 96 nodes over 21 days (about 392,000 GPU hours) on infrastructure in Germany and Finland; the model card estimates energy consumption at about 950 MWh including data centre overhead. Post-training combines supervised fine-tuning on curated German and English data with reinforcement learning, among others in retrieval and code environments.

The UniBPE tokenizer with 128,000 entries is designed bilingually. In its blog, Aleph Alpha reports about 4.9 bytes per token on German web text (FineWeb-2) and 4.58 bytes per token on English (FineWeb). For German compounds, that means less fragmentation and therefore fewer tokens per document than with English-centric tokenizers; the actual saving depends on the corpus.

Benchmarks: Aleph Alpha’s figures

In the tech report (Tables 28 and 29), Aleph Alpha compares Kolibri-1 with open MoE models of the same class, i.e. roughly 3 to 6 billion active parameters. These are vendor figures; all models were measured by Aleph Alpha in the same environment.

BenchmarkKolibri-1Qwen3.5 35B-A3BQwen3.6 35B-A3BGemma 4 26B-A4BNemotron 3 Nano 30B-A3B
Overall (EN)75.574.771.471.965.6
Overall (DE)70.869.867.366.359.3
GPQA Diamond (DE)81.384.280.680.149.6
MMLU-ProX CoT (DE)75.581.781.981.161.0
AIME 2025 (DE)87.576.782.988.184.4
LiveCodeBench v685.977.882.582.371.3
SWE-Bench Verified66.471.673.857.838.6
Tau2-Bench Telecom94.797.799.145.345.9
BFCL v4 (overall)61.470.567.268.261.5
BFCL v3 (multi-turn)39.854.053.553.447.9
RGB Closed-Book51.081.079.079.080.0
AA-Omniscience non-hallucination44.011.156.714.319.0
Industry RAG (DE)67.570.065.849.446.1

Three observations. First, on the overall scores Kolibri-1 sits just ahead of the comparison models in its class in both languages, with the largest margin in German. Second, mathematics, code and retrieval-grounded tasks are its strengths: AIME 2025 in English 96.9, Honeypot 80.8, Industry RAG in English 89.7. Third, the results for closed-book factual knowledge (RGB Closed-Book 51.0) and multi-turn tool calling (BFCL v3 multi-turn 39.8, TerminalBench 2.1 27.7) show clear gaps; Aleph Alpha itself positions the model for use with retrieval and describes it as abstaining rather than guessing when the context does not support an answer. The high non-hallucination rate of 44.0 on AA-Omniscience fits this design: the model answers incorrectly less often because it more often does not answer at all.

For long context, the model card reports RULER scores for the base model of 63.2 at 1 million tokens (Qwen3.5 35B: 57.5) but 69.8 at 256k versus 80.1 for Qwen; for efficiency, the model card recommends inputs of up to 262,144 tokens.

Operations: hardware and vLLM

Kolibri-1 is designed for self-hosting. The FP8 weights take about 78 GB; embeddings, LM head, norms and MoE router are in BF16, the KV cache in FP8.

GPUs
Minimum (model card)2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300
Recommended (model card)2× H100 SXM5, 2× H200, 1× B200 or 1× B300
Concurrent 256k requests in FP8 (tech report, Table 31)2× H100: ~18, 4× H100: ~67, 1× B200: ~31, 2× B300: ~157

Serving starts via vLLM with Aleph Alpha’s plugin:

pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

Alternatively there is a container image at ghcr.io/aleph-alpha/aleph-alpha-inference. The model card recommends temperature 1.0, top-p 0.97 and top-k 128. The server is OpenAI-compatible, so Kolibri-1 can be run like other self-hosted models as a backend for CompanyGPT or behind the AI Gateway. A hosted Aleph Alpha API for Kolibri-1 is not currently documented; the model card points to sales for deployment and specialisation. WirtschaftsWoche and other media report planned offerings on STACKIT together with the Schwarz Group.

Data protection, regulation and sovereignty

In the tech report, Aleph Alpha describes a development pipeline that takes the EU AI Act, the General-Purpose AI Code of Practice and the GDPR into account from the start: the data pipeline filters against illegal, harmful and pirated content and redacts personal data from sources; the model is meant to abstain when the context does not support an answer. The model card also lists limitations (systemic biases from the training data, knowledge outdated after June 2026, hallucinations) and recommends prompt design, application-level validation, output filters and red-teaming; the model is not recommended for unsupervised high-stakes decisions.

For the sovereignty question, the deployment path matters most: with open weights on your own hardware, neither prompts nor documents leave your environment. That distinguishes Kolibri-1 from hosted models such as OpenAI GPT or Anthropic Claude, where EU data residency is established through cloud deployments such as Microsoft Foundry or Bedrock. Training itself took place in Germany and Finland according to Aleph Alpha.

Positioning: Kolibri-1, Pharia, Mistral and OpenEuroLLM

Kolibri-1Pharia-1 LLM-7BMistral (open models)OpenEuroLLM
VendorAleph Alpha (Heidelberg)Aleph AlphaMistral AI (Paris)EU consortium
ArchitectureMoE 78.1B / 3.46B activedense 7Bdense and MoE, depending on modelin development
Context1M tokens8k tokensdepending on model–
LanguagesGerman, Englishmultilingual (7 EU languages)multilingualall EU languages (goal)
LicenceApache 2.0Open Aleph License (non-commercial)Apache 2.0 (open models)open (goal)
Hosted APInot documentedPhariaAIMistral La Plateforme, Azure, AWS and others–

Kolibri-1 is the model for organisations that need German as their main language, a long context and operation on their own hardware, and can do without image input and additional languages for that. Anyone looking for broad language coverage or a hosted API is better served by Mistral; anyone waiting for a consortium-backed EU solution keeps an eye on OpenEuroLLM.

Our recommendation

For public authorities, industrial companies and regulated sectors in the DACH region, Kolibri-1 is a candidate for the shortlist of self-hosted models: open licence, German language quality within its class, 1 million tokens of context, reasoning levels and a deployment path via vLLM that gets by with two H100s or one B200. Check three points before deploying: quality in multi-turn tool calling (where the model trails its comparison models), the handling of factual questions without retrieval (the model is designed for RAG) and the operating cost of the GPU infrastructure compared with an EU cloud deployment of a hosted model. For the existing Aleph Alpha portfolio: Kolibri-1 replaces Pharia-1 LLM-7B as our recommendation.

We evaluate Kolibri-1 on a per-project basis as a backend for CompanyGPT and as a model behind the AI Gateway. For an assessment of whether Kolibri-1 fits your requirements, contact innFactory AI Consulting.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Frequently Asked Questions

What is Aleph Alpha Kolibri-1?

Kolibri-1 is a language model for German and English released by Aleph Alpha Research on 3 October 2026. It is a Mixture-of-Experts transformer with 78.1 billion parameters, 3.46 billion of which are active per token, processes contexts of up to 1,048,576 tokens, supports reasoning at four effort levels (none, low, medium, high) plus tool calling, and is available under Apache 2.0 on Hugging Face. According to Aleph Alpha it was developed end to end in Germany and trained on infrastructure in Germany and Finland, with the EU AI Act, the GPAI Code of Practice and the GDPR in mind.

What hardware does Kolibri-1 need?

The FP8 weights take about 78 GB. The model card lists two A100 80 GB, two H100 SXM5, one H200, one B200 or one B300 as the minimum and recommends two H100 SXM5, two H200, one B200 or one B300. Serving runs on vLLM with the aleph-alpha-inference package, which ships reasoning and tool-call parsers for Kolibri. According to the tech report, two H100s fit about 18 concurrent 256k-token requests in FP8, one B200 about 31.

How good is Kolibri-1 in German?

In its tech report Aleph Alpha reports an overall score of 70.8 on German benchmarks (English: 75.5). Among MoE models with roughly 3 billion active parameters, that places Kolibri ahead of Qwen3.5 35B-A3B (69.8), Qwen3.6 35B-A3B (67.3), Gemma 4 26B-A4B (66.3) and Nemotron 3 Nano 30B-A3B (59.3). Individual results: GPQA Diamond in German 81.3, MMLU-ProX 75.5, AIME 2025 in German 87.5. The figures come from the vendor; independent measurements were not available at launch.

Can Kolibri-1 be used in a GDPR-compliant way?

Kolibri-1 is built to run on your own infrastructure: the Apache 2.0 weights can be served with vLLM in your own data centre or at an EU cloud provider, so no data flows to a model vendor. According to Aleph Alpha, the data pipeline filters illegal, harmful and pirated content and redacts personal data from training sources. A hosted Aleph Alpha API for Kolibri-1 is not currently documented; media reports mention planned offerings on STACKIT together with the Schwarz Group.

How does Kolibri-1 differ from Pharia and Luminous?

Luminous and Pharia-1 were Aleph Alpha's earlier model generations; Pharia-1 LLM-7B is a dense 7-billion-parameter model with 8k context under the non-commercial Open Aleph License. Kolibri-1 is a new development: Mixture-of-Experts with 78.1 billion total parameters, 1 million tokens of context, reasoning levels, tool calling and a new UniBPE tokenizer with 128,000 entries. The model card makes no reference to Pharia or the T-Free architecture; according to the tech report, Kolibri was trained from scratch.

What does the Cohere merger mean for Kolibri-1?

Aleph Alpha and Cohere signed a definitive merger agreement on 16 September 2026; per Cohere, closing is still pending. Kolibri-1 is released under the Aleph Alpha Research name with an open Apache 2.0 licence that applies regardless of corporate structure. According to media reports, offerings based on Kolibri are to become part of the joint sovereign AI offering on STACKIT. Neither company has published a roadmap for further Kolibri sizes or for combining it with Cohere's Command models.

Consultation for this model?

We help you select and integrate the right AI model for your use case.