senn-techsenn-tech
AI & Development
AI & Development2026-10-05· By Franz Senn

Aleph Alpha's Kolibri: a German tokenizer that will not fit our GPUs

On 3 October 2026 Aleph Alpha published the weights of Kolibri, followed by a press release two days later. It is a mixture of experts with 78.1 billion weights in total, 3.46 billion active per token, 262,144 tokens of context natively and a little over a million by extrapolation, and a tokenizer built for German rather than adapted to it. The licence is Apache 2.0 without an additional clause, the repository carries the unmodified standard text.

What pushed us to look was not the announcement but the argument around it. A widely shared post presented Kolibri as a German model on par with considerably larger open models, adding in one line that the benchmark figures came from the vendor. That is the point where we stop reading and start measuring. Three questions matter to us: what the primary source actually claims, what survives contact with our own building, and what running it here would cost.

What the model is, technically

Every figure below comes from the model card or the 189 page technical report, both pulled directly from Aleph Alpha.

PropertyValue
Parameters, total78,103,074,560
Parameters active per token3,457,573,120
Layers50, sliding window to full attention in a 4:1 ratio, window size 512
Experts384 per layer, one shared and six active, sigmoid routing without renormalisation
Weight precisionFP8 e4m3 with 128×128 block scales, embeddings and router in BF16
Context262,144 native, 1,048,576 through YaRN extrapolation
Weight sizeroughly 78 GB, vendor minimum two A100 80 GB
Tokenizer128,000 entries, new method called UniBPE
Pre-training data20 trillion tokens: 62.5 percent English, 23.9 percent German, 13.6 percent code, about 24 percent synthetic
Compute768 B200 GPUs, 21 days, 392,000 GPU hours for pre-training alone

Two details from the appendix are worth carrying over. The synthetic data was generated with GLM-5.2, GLM-5.3 and Qwen3.8-27B as teacher models, and the German web corpus was grown past two trillion tokens with a pipeline of their own. And the training location is a single sentence in 189 pages: teams in Germany developed the model end to end and trained it on infrastructure in Germany and Finland. Which Finnish cluster is never stated, and the press release only says Europe.

The tokenizer is the part we can measure ourselves

A claim about a model cannot be checked without the model. A claim about a tokenizer can, so we lined up two files: the tokenizer.json from the Kolibri repository and the tokenizer.json of the model that actually serves our traffic, Qwen3.8-Flash-Next-Hybrid on our 4× RTX 5090 host. Vocabulary 127,900 against 248,044 entries.

Start with the single word from the discussion, Bundesverfassungsgericht:

TokenizerTokens for that one word
Kolibri UniBPE2
our production lane4
o200k, the ChatGPT lineage6

So the figure circulating on LinkedIn reproduces, and our own model sits in between. A real text is more interesting. We took four paragraphs we write for cases like this: a forwarding confirmation, a data protection clause for our portal, a dangerous goods declaration, a demand note about a wrong quantity. 1,526 UTF-8 bytes, one measurement run per tokenizer, corpus and script kept in the dossier for this post.

Text typeKolibriour laneo200kKolibri against our lane
Business and freight German, 4 texts295360341−18.1 %
English control texts, 2 texts112114112−1.8 %
our own blog corpus, 14 posts5,7306,2746,010−8.7 %
same German text, umlauts spelled ae, oe, ue354399376−11.3 %

Bytes per token on the business German: 5.17 for Kolibri, 4.24 for our lane, 4.48 for o200k. The vendor's figure for German web text is 4.90, so our formal prose sits above it, which is what long compounds do. The English control shows there is no penalty to pay: 112 against 114 tokens.

Two caveats belong to that number, and neither is in the model card.

Umlauts carry part of the gain. Same four paragraphs, once with correct spelling and once with ae, oe, ue and ss, which is exactly what exports from older ERP systems hand us: Kolibri drops from 5.17 to 4.31 bytes per token and the gap to our lane narrows from 18.1 to 11.3 percent. Kolibri spends its strongest trick on correctly encoded umlauts. Anyone piping German text with broken encoding into an AI pipeline wastes money with every tokenizer, and with this one most of all. Fixing the encoding is the cheaper lever than swapping the model.

Technical mixed text eats the advantage. We also measured our own 14 most recent German posts, 25,505 bytes of real mixed documentation. There the difference is only 8.7 percent, spread 5.9 to 12.6 percent, median 8.8. The word counts we recorded alongside explain it: 149 to 180 English technical terms per text against 11 to 33 non-ASCII characters. Kolibri is a strong German tokenizer, not a conjurer.

The quality table stays a vendor table

We looked for an independent number in two places on 5 October and found none. The Artificial Analysis sitemap returns 200 and contains zero Kolibri URLs. OpenRouter lists 464 models with zero hits. Outside of Aleph Alpha there was no measurement to read.

The chart being passed around is figure 46 of the technical report. Its horizontal axis is neither memory nor throughput but active parameters. On that axis a model with 3.46 billion active parameters lands far left by construction and lands on the hull. The report says this openly: Qwen3.8-27B reaches the highest aggregate score, but it is dense and carries almost eight times the active parameters. Exclude dense models from the axis and you get the hull you expected. The same table shows how loosely such comparisons read: within one measurement run Qwen3.6 35B-A3B scores 71.4 on the English aggregate, below Qwen3.5 35B-A3B at 74.7. That is not evidence against Kolibri, it is a reminder that the baselines are snapshots of different versions.

Then there are the rows the vendor itself labels weak, and they are the ones that touch our work:

EvaluationKolibribest comparison in the same table
German aggregate70.879.9, Qwen3.8 27B
SWE-Bench Verified66.473.8, Qwen3.6 35B-A3B
TerminalBench 2.127.739.7, Qwen3.5 35B-A3B
RGB closed book51.093.0, Nemotron 3 Super
HELMET at 8k context83.795.3, Qwen3.5 35B-A3B
AA-Omniscience index−32.8−15.3, Qwen3.6 35B-A3B

That last row is the one to think about. Kolibri was demonstrably trained to decline when the context does not carry the answer, with purpose built reinforcement environments and abstention data. On the axis Aleph Alpha chose for exactly that property, the model scores worse than the smaller Qwen3.6 35B-A3B: −32.8 against −15.3, with a truthful-answer rate of 44.0 against 56.7. The method is documented in detail, the result on that very method is the weakest of the comparison group. That is the difference between a trained intention and a measured outcome.

The hardware arithmetic for our building

We looked at both GPU hosts on the evening of 5 October instead of calculating.

Hostmeasured stateKolibri
.180.3, 4× RTX 5090cards at 32,061, 31,947, 31,947 and 31,947 MiB of 32,607 MiB, vLLM build 0.1.dev20073only by evicting the production lane, and the version pin is not met
.180.202, 1× RTX 509030,415 MiB used, vLLM 0.30.0, 184 GB RAM, EPYC 750278 GB of weights against a 32 GiB card, no
.180.211embeddings, reranking and document processing, no language lanenot a candidate for a 78 GB model at all

The serving path is the real constraint. The package aleph-alpha-inference, which is what makes the architecture loadable in vLLM at all, states in its project file that each release supports exactly one vLLM minor version, here vllm>=0.29.0,<0.30.0. Our production lane runs a development build of our own, the standby host runs 0.30.0. Kolibri would therefore be a second container with a second vLLM, not another model folder in an existing image. The package repository was created on 28 September, has 19 stars and two open issues.

The llama.cpp route several commenters mention exists, but with a patch. A community repository ships Q4_K_M at 47.5 GB, Q6_K at 64.2 GB and Q8_0 at 83.1 GB, plus a patch that applies to one specific llama.cpp commit. The reason is in the same README: the router selects on logits plus expert bias and weights with the unrenormalised sigmoid, while llama.cpp's existing DeepSeek V3 routing treatment computes this differently. The author of that conversion measures 11.9 tokens per second for Q4_K_M in CPU-only operation at 46.6 GB of memory, and the CUDA path falls back to MoE execution without kernel fusion. On our standby host, an EPYC 7502 without AVX-512, that is not a language lane but a queue.

What we make of it

No installation, no test lane, revisit. Four conditions make us measure again: an independent evaluation that runs against Qwen3.8-Flash-Next rather than a mean of open baselines. A supported vLLM version that matches one of our hosts. A free machine with two 80 GB cards or an H200. And the case where contracts, consignment notes and ERP text really do move into a retrieval pipeline at this company, because then UniBPE becomes an interesting reference point for German token costs whether or not the weights ever run here.

What stays is the working conclusion that for a German language model the productive question is not the benchmark row but the tokenizer and the memory bandwidth. Both can be tested without the model, and we now have numbers for both.

Further reading

Questions?
Could Kolibri run on our existing AI infrastructure?+

Not as an additional lane. The FP8 weights occupy roughly 78 GB, so the only candidate is our 4x RTX 5090 host with 127 GiB of card memory, and its four cards are sitting at 31.9 to 32.1 GiB of 32.6 GiB each because the production language lane lives there. The version pin settles it: the serving package aleph-alpha-inference requires vllm>=0.29.0,<0.30.0, while our standby host runs vLLM 0.30.0 and the production host a custom development build.

Is Kolibri better than the model we run today?+

No source answers that question. Aleph Alpha's own evaluation table does not contain Qwen3.8-Flash-Next, which is the model carrying our requests, and in that same table the dense Qwen3.8 27B scores 79.9 on the German aggregate against Kolibri's 70.8. A blind test mentioned in a LinkedIn comment claims the opposite on five prompts without a protocol. As of 5 October there is no independent measurement of this model anywhere we could find.

Is the German tokenizer worth anything to us?+

On formal German yes, on technical mixed text barely. Across four business and freight texts we wrote ourselves, Kolibri needs 18.1 percent fewer tokens than our production lane. Across our own 14 published German posts it is only 8.7 percent. And when the umlauts are transliterated to ae, oe and ue, which is what legacy ERP exports deliver, the advantage shrinks from 18.1 to 11.3 percent.