Sunday, Oct 4 | --:--
Back to home
Models Open weights Europe

Aleph Alpha Releases Kolibri: Apache 2.0 German/English MoE, 3.46B Active

Aleph Alpha’s October 3 release, Kolibri, is an English-German mixture-of-experts model with 78.1 billion total parameters and 3.46 billion active, context up to about 1 million tokens (the model card recommends serving at or below 262,144), and full weights on Hugging Face under Apache 2.0. The company says it was trained in Germany and Finland and that 21.3% of pre-training tokens are German. Scores such as AIME 2025 English 96.9 are Aleph Alpha’s own harness at high reasoning effort, not an Arena or Artificial Analysis placement.

Times of AI Desk 6 min read Germany View as Markdown
Cover illustration for Aleph Alpha Releases Kolibri: Apache 2.0 German/English MoE, 3.46B Active

A European lab just put a bilingual open-weight model on Hugging Face under Apache 2.0. The capability claim is still the lab’s own scoreboard.

Aleph Alpha, in an October 3, 2026 post dated to the Day of German Reunification, says Kolibri is an English-German mixture-of-experts transformer with 78 billion total parameters and about 3 billion active, context up to 1 million tokens, downloadable in full on Hugging Face under Apache 2.0. The Hugging Face card for Aleph-Alpha/Kolibri-1 is more exact: 78,103,074,560 total parameters, 3,457,573,120 active per token (3.46 billion), release date October 3, 2026, license Apache 2.0, knowledge cutoff June 18, 2026 for both English and German. The card says quality was validated up to 1,048,576 tokens and recommends at most 262,144 for serving efficiency and complex tasks. Weights are FP8, with a stated memory footprint of about 78 GB.

What actually shipped

Kolibri, as published
Repo Aleph-Alpha/Kolibri-1
Total / active 78.1B / 3.46B (card); blog prose says ~3B active
Context Up to 1,048,576; card recommends ≤262,144
License Apache 2.0 on the weights
Reasoning none, low, medium, high; tool calling
Knowledge cutoff EN and DE, June 18, 2026
German data Blog: 21.3% of pre-training tokens
Where trained Blog: Germany and Finland, under European law
Pre-training compute Card: 768 NVIDIA B200s, 21 days (pre-training only)

The blog’s build table matches the card on the 3.46B active figure, four reasoning efforts, and a native long-context train at 262,144 tokens, with the million-token length treated as an extension. Serving is not stock vLLM alone: both the post and the card require Aleph Alpha’s aleph-alpha-inference plugin (vllm serve Aleph-Alpha/Kolibri-1 with Kolibri reasoning and tool parsers).

Aleph Alpha’s sovereignty pitch is deployment control plus a German/English data mix, not a new legal status. The blog says the model is aimed at public administration, industrials, and aerospace, that customers can run it on-premise, and that 21.3% of pre-training tokens are German (about 4.3 trillion of a 20 trillion token run). The card’s training-data summary describes the filtered bilingual corpus with a German share nearer 24%; that is a mix description, not a second measured benchmark. This piece uses 21.3% for the share of tokens the blog says the run actually saw.

Claims vs checks

Aleph Alpha says benchmarks used its own harnesses and, where applicable, each model’s highest reasoning effort. On that table, Kolibri’s AIME 2025 is 96.9 English and 87.5 German; GPQA Diamond English is 84.3; LiveCodeBench v6 is 85.9. The same table’s AA-Omniscience Index (public set) for Kolibri is −32.8 — a lab number, not a placement on Artificial Analysis. The blog’s line that Kolibri “matches models with up to four times its active parameter count, such as Nemotron 3 Super” is that comparison, not an independent audit.

The card’s eval note says category averages are unweighted means and that Kolibri was run at reasoning effort high. Those settings are not a license to rank Kolibri against closed frontier models on Arena or Artificial Analysis.

A public-board check on the evening of October 4, 2026, did not show Kolibri in the visible Artificial Analysis Intelligence Index table or on the LM Arena page opened then. That is a snapshot of what was on screen, not a search of every alias, and it is not a quality verdict.

Limits

  • All benchmark figures above are Aleph Alpha harness results. They were not re-run here.
  • Blog prose says “~3B active”; the card and the blog’s own spec table say 3.46B. The exact card figure is the one used for size.
  • German-share wording differs between the blog’s 21.3% of pre-training tokens and the card’s corpus-mix description. They are not averaged.
  • “Pareto frontier” for quality versus serving cost is the company’s framing.
  • Arena / Artificial Analysis: not observed on the October 4 evening snapshot. That check was not refreshed later the same evening. No Elo is inferred.
  • Weights were not downloaded; license, size, and release date are what the blog and model card state.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading