Aleph Alpha launched Kolibri 1 on October 3 — a 78-billion-parameter open-weight model built and trained entirely in Europe, released under Apache 2.0, and deliberately designed with no hosted API. You self-host it or you don’t use it. For the growing number of organizations scrambling to meet EU AI Act enforcement requirements that became active in August 2026, that constraint isn’t a limitation. It’s the entire selling point.
78 Billion Parameters, 3.46 Billion Active
Kolibri uses a Mixture-of-Experts architecture: 384 experts per transformer layer, with only 6 firing per forward pass. The result is 78.1B total parameters but just 3.46B active per token — a 22:1 sparsity ratio that makes inference costs comparable to a 3-4B dense model. A community member already confirmed it runs on a single RTX Pro 6000 in FP8 at roughly 170 tokens per second. You don’t need a cluster to experiment.
Benchmarks tell a respectable story: 96.9% on AIME 2025 math, 85.9% on LiveCodeBench v6, 84.3% on GPQA Diamond. That puts Kolibri in range of models with four times more active compute. The context window extends to one million tokens, with 262K being the recommended production ceiling for reasonable latency. Official minimum deployment is two H100 SXM5s — the single-GPU path is community-validated but not officially supported.
However, it’s not the best model on every benchmark. Tool calling scores 61.4 on BFCL v4 versus Qwen 3.6’s 67.2, and software engineering on SWE-Bench comes in at 66.4 versus 73.8. Qwen 3.8 at 27B active parameters actually outscores Kolibri on general English benchmarks. Aleph Alpha isn’t positioning this as the universal winner — they’re positioning it as the best model you can fully control.
Why “No API Endpoint” Is the Right Design for Kolibri
Aleph Alpha built Kolibri with no hosted inference service. Every organization that uses it must run it on their own infrastructure. This sounds like a gap, but the regulatory environment makes it a feature. EU AI Act high-risk AI provisions are now enforceable. For AI deployed in healthcare, government, aerospace, and finance, that means conformity assessments, risk management documentation, and demonstrable data governance. Routing sensitive data through a third-party API creates GDPR cross-border transfer liability that self-hosting simply eliminates.
The penalties aren’t hypothetical: up to €35 million or 7% of global annual turnover for serious EU AI Act data sovereignty violations. Organizations in regulated industries can no longer treat “which model performs best on benchmarks” as the primary selection criterion. They need a model they can audit, a training pipeline they can verify, and infrastructure they fully control. Kolibri ships with a 189-page technical report and full supply-chain transparency documentation — precisely the kind of compliance artifact those industries require.
Related: FTC Probes AI Agent Safety: What Developers Must Know
Getting Kolibri Running
Kolibri requires Aleph Alpha’s custom vLLM plugin — stock vLLM won’t work. Install the aleph-alpha-inference package first, then serve via vLLM. Once running, the model exposes an OpenAI-compatible API surface, so existing chat-completion code connects without modification.
pip install "aleph-alpha-inference>=1.0"
vllm serve Aleph-Alpha/Kolibri-1 \
--kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
Kolibri model weights on Hugging Face are available under Apache 2.0, approximately 78GB in FP8 format. Recommended sampling: temperature 1.0, top_p 0.97, top_k 128. The model supports four reasoning effort levels — none, low, medium, and high — which let you trade response time for thoroughness at query time. Use effort=none for classification or extraction tasks to avoid the model reasoning through problems that don’t need it. For a more detailed setup walkthrough, see the Kolibri deployment guide.
The Signal Beyond the Benchmarks
Kolibri landed as the top story on Hacker News on October 4 with 498 points and 298 comments — unusually high for a European AI company. Most North American developers don’t closely follow Aleph Alpha. The engagement reflects genuine interest in what a credible European open-weight model looks like, not just Germany’s public sector procurement story.
The model is bilingual English-German, with 62.5% English, 23.9% German, and 13.6% code in its training data. Aleph Alpha curated 80% of the German-language training data in-house because sufficient public German data simply doesn’t exist at the quality needed. That investment in German data quality extends the sovereignty pitch beyond hosting location — it covers training-data provenance as well. Read the full Kolibri 1 announcement for the complete technical breakdown, including the 189-page report.
European AI has had a credibility problem: strong regulatory ambition, weaker model execution. Kolibri doesn’t close that gap entirely, but it demonstrates that a European lab can train a competitive open-weight MoE at scale on European infrastructure. For regulated industries globally — not just in Germany — that’s a meaningful option that didn’t exist a week ago.
Key Takeaways
- Kolibri 1 launched October 3 — 78B total parameters, 3.46B active per token, Apache 2.0, no hosted API by design.
- EU AI Act high-risk enforcement is active as of August 2026. Self-hosting is no longer just a preference for regulated industries — it’s a compliance strategy.
- Benchmarks are competitive (96.9% AIME 2025, 85.9% LiveCodeBench) but not universal leaders — tool calling and SWE-Bench are weaker than Qwen 3.6.
- Deployment requires the aleph-alpha-inference vLLM plugin; the API surface is OpenAI-compatible once running.
- The 498-point HN reception signals genuine developer interest in sovereign, auditable AI — Kolibri’s audience extends well beyond Germany.













