Builders can download Aleph Alpha’s Kolibri weights under Apache 2.0 and run the German-English LLM on their own infrastructure. The documented serving setup requires Aleph Alpha’s vLLM integration, and the company recommends a 262,144-token operating length despite the advertised one-million-token maximum.
Kolibri also needs more hardware than its active parameter count might suggest. Its mixture-of-experts architecture activates 3.46 billion parameters per token, but the complete model contains 78.1 billion parameters. Sparse computation reduces the work performed during generation; it doesn't make the other weights disappear.
In its October 3, 2026 announcement, Aleph Alpha positions Kolibri for public administration and industrial workloads where German-language capability and deployment control matter. Builders get a commercially permissive, self-hostable model with reasoning and tool calling. Its performance claims still require independent scrutiny.
Apache-2.0 Weights Make Self-Hosting a Real Option
The Kolibri-1 release provides downloadable weights under Apache 2.0. That commercially permissive license lets organizations use, modify, and redistribute the model subject to its terms, without relying exclusively on a vendor-operated inference API.
For teams that need to keep inference inside infrastructure they control, downloadable weights make that choice possible. They still need suitable hardware, operational expertise, and application-level safeguards.
German is a deliberate part of the model’s design. Aleph Alpha says it developed a bilingual German-English tokenizer and that German accounted for 21.3% of pre-training tokens. It also describes using organic German-language material rather than relying primarily on translated English data.
German-speaking teams therefore have a concrete reason to evaluate Kolibri. Whether it handles particular administrative terms, manufacturing procedures, or legal documents correctly remains a workload-specific question.
Aleph Alpha describes Kolibri as a sovereign model for regulated environments. That is its deployment and product positioning, not a blanket compliance guarantee. A permissive weight license and local inference provide control; they don't independently establish that an application satisfies privacy law, sector requirements, or its organization’s approval process.
The vLLM Integration Is Part of the Deployment
Kolibri’s documented serving route uses Aleph Alpha’s inference package with vLLM, including dedicated reasoning and tool-call parsers. Builders should treat that integration as a requirement for the published setup. They shouldn't assume an unmodified vLLM installation is sufficient.
The parsers do different jobs. A reasoning parser separates the model’s reasoning output from its final response. A tool-call parser turns the model’s generated function-call format into structured calls that application code can process. Correct parsing matters even when the generated text looks plausible.
The supplied model card’s chat template also distinguishes reasoning, assistant responses, tool calls, and tool responses. Multi-step workflows need that structure preserved: a successful single-turn text generation test doesn't establish that an agent loop works correctly.
With four reasoning-effort levels, Kolibri gives applications control over how much reasoning the model performs, according to Aleph Alpha. This offers a tuning option for workloads that mix straightforward extraction with more demanding analysis. The resulting quality, latency, and output-length trade-offs need measurement.
For an initial integration, builders should validate three paths separately:
- Ordinary responses, including the separation of reasoning from the final answer.
- Structured tool calls, including argument validation and tool-response handling.
- Reasoning-effort settings, compared on representative tasks.
Tool calling doesn't provide permission to execute an action. The application still needs to enforce access controls and decide which operations require approval.
The Practical Context Target Is 262K, Not 1M
Kolibri’s maximum context configuration is 1,048,576 tokens. Its trained sequence length and recommended operating length are 262,144 tokens, commonly described as 256K.
Aleph Alpha recommends the shorter length for serving efficiency and complex tasks. Serving beyond it requires explicit vLLM overrides, so the larger configuration is an option to evaluate, not the default operating target implied by the headline specification.
Three separate questions matter here:
- How much input can the server accept?
- How well does the model use information across that input?
- What memory, latency, and concurrency costs does processing it impose?
A maximum context setting answers only the first question. It doesn't demonstrate reliable retrieval, synthesis, or reasoning throughout a million-token input.
The official model specifications explicitly list both the maximum and recommended lengths. For an initial production evaluation, 262,144 tokens is the more defensible starting point. Teams should expand beyond it only when a workload benefits and their measurements support the additional cost.
Even at the recommended length, longer prompts aren't automatically better. Retrieval-augmented generation may let an application select relevant passages instead of repeatedly processing an entire collection. The useful comparison is whether additional context improves answers enough to justify its operational cost.
The 3.46B Active Count Doesn't Describe Memory Requirements
The model card gives Kolibri’s precise size as 78,103,074,560 total parameters, with 3,457,573,120 active per token.
In a mixture-of-experts model, routing selects a subset of expert networks for each token. A large collection of learned weights can thus contribute across different inputs without every expert being used during every generation step. This lowers active computation, while the full set of weights still needs to be stored and made available to the serving system.
Kolibri-1 is distributed with primarily FP8 weights. Components including embeddings, the language-model head, normalization layers, and the router remain in bfloat16. Aleph Alpha lists an approximately 78 GB weight memory footprint.
Its published minimum configurations include a single H200 or two A100 80 GB GPUs. Recommended configurations include two H100 SXM5 GPUs or two H200 GPUs, among other options. These are vendor-listed requirements, not independently measured capacity guarantees for every workload.
Beyond the weights, serving requires working memory and a key-value cache, which holds attention state for processed tokens. Context length and concurrent requests therefore affect whether a deployment fits and performs acceptably.
Kolibri is computationally sparse, but it isn't a 3.46-billion-parameter model in storage terms. Hardware planning needs to account for both its total size and its active computation.
Vendor Benchmarks Show Strengths, Not a Universal Win
Aleph Alpha publishes comparisons covering mathematics, knowledge, coding, tool use, and long-context tasks. It also claims a favorable quality-versus-serving-cost position.
These are vendor-reported results. The supplied release evidence does not establish independent reproduction, so the scores should guide evaluation without settling purchasing or deployment decisions.
A selection from the company’s published comparison shows mixed results:
| Benchmark | Kolibri | Qwen3.6-35B-A3B |
|---|---|---|
| AIME 2026, German | 90.0 | 84.4 |
| GPQA Diamond, German | 81.3 | 80.6 |
| BFCL v4 overall |
Frequently Asked Questions
4 questions
1Can Kolibri be used commercially?
Yes. Kolibri’s downloadable weights are released under Apache 2.0, which allows use, modification, and redistribution subject to its terms. Organizations still need to assess their applications separately: the weight license does not establish privacy compliance, sector approval, or the safety of an automated workflow.
2
Sources
- October 3, 2026 announcementaleph-alpha.com
- Kolibri-1 releasehuggingface.co
- official model specificationsaleph-alpha.com





