A 78 B model on one 96 GB card, 21% faster under load than the FP8 release.
Mixed NVFP4/FP8 build of Aleph-Alpha/Kolibri-1-BF16 at
45.81 GB, on one 96 GB card.
All 50 layers of routed experts at 4 bits; attention and shared experts are the FP8 weights Aleph Alpha
ships; embeddings, lm_head, routers and norms stay BF16.
Two builds, one measurement. We quantized the experts two ways and ran both against Aleph Alpha's own FP8 release under one protocol. This build keeps the release's FP8 attention and is the faster. The NVFP4 build keeps attention in BF16 and is a little more accurate on knowledge. On tool calling all three are one band.
Why this quant
- 🗜️ 45.81 GB against 78.83 GB for the FP8 release. The release takes 73.55 GiB of a 96 GB card for weights; this build takes 42.78 GiB. The KV pool at 32K context grows from 410,315 tokens to 1,461,977, 3.6× as many.
- ⚡ 1,195 tok/s at concurrency 32, against 986 for the release (+21%). One user at a time it goes the other way: 142.6 tok/s against 159.1.
- 🎯 Tool calling is level with the release. 77.9 over five runs against 78.5 over five; the run-to-run spread is about 1.1 on both sides.
- 📉 Knowledge is about a point lower. 85.0 against 86.1. Both of this build's runs sit below both of the release's, and a gap of 1.1 is at the edge of what a 1,170-item suite resolves, so we report it as measured and do not call it a tie.
- 🧩 Only the experts changed. Attention and shared-expert weights are Aleph Alpha's FP8 bytes, copied unmodified, so none of the difference above comes from re-quantizing them.
Serve it
hf download primitive-ai/Kolibri-1-mixed-NVFP4-FP8 --local-dir ./Kolibri-1-mixed-NVFP4-FP8
docker run --gpus all --ipc=host -p 8000:8000 -v $PWD:/models \
ghcr.io/aleph-alpha/aleph-alpha-inference:1.0.0-vllm0.29.0 \
/models/Kolibri-1-mixed-NVFP4-FP8 --served-model-name kolibri \
--kv-cache-dtype fp8 --max-model-len 32768 --gpu-memory-utilization 0.92 \
--reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice
Stock vLLM does not know the Kolibri1ForCausalLM architecture. Aleph Alpha's plugin image adds it
along with the kolibri1 parsers, and nothing here is patched on top of it. The image is vLLM
0.29.0 with aleph-alpha-inference 1.0.0 (sha256:9a56ab1691f8bc8eb0fc82f0f534bfdd36da2218fa79b01b83448739fbda403b). Boot takes about 130 s on an idle card,
and the NVFP4 experts run on FlashInfer CUTLASS.
Aleph Alpha recommends temperature 1.0, top_p 0.97, top_k 128. Our numbers use the fixed
protocol above for every build, so they are comparable with each other, not with Aleph Alpha's own.
The measurement servers ran without the two parsers, because the tool-calling suite does not pass
tools=. Thinking is on by default; send enable_thinking: false to turn it off.
Measured
One RTX PRO 6000 Blackwell, 96 GB, one card. The 1,170-item knowledge suite and the 200-item
tool-calling suite, temperature 0.6 / top_p 0.95 / top_k 20, thinking on, a 16,384-token
budget, concurrency 32, auto-scored with no LLM judge. FP8 KV cache for every row, which is how
Aleph Alpha evaluates the release. Throughput is 8K in / 512 out, prefix-cache free, two seeds per
cell. Every row ran on the same box on the same day.
| build | size | knowledge | tool-calling | call | abstain | finished | tok/s @1 | tok/s @32 | KV pool |
|---|---|---|---|---|---|---|---|---|---|
| Kolibri-1, FP8 release | 78.83 GB | 86.1 | 78.5 | 86.1 | 48.0 | 97.0% | 159.1 | 986 | 410,315 |
| NVFP4 | 47.70 GB | 86.2 | 77.5 | 84.9 | 48.0 | 96.4% | 124.4 | 1076 | 1,400,826 |
| this repo, mixed | 45.81 GB | 85.0 | 77.9 | 85.5 | 47.5 | 96.6% | 142.6 | 1195 | 1,461,977 |
Tool-calling is the mean of five runs per build, with a within-build standard deviation of 0.6
to 1.2, so that column is one band. Knowledge is two runs per build (86.0/86.2, 85.9/86.5, 84.8/85.1 strict, in table order).
On knowledge the release and NVFP4 are 0.1 apart; the mixed build is 1.1 below the release and 1.2 below NVFP4, on both of its runs. KV pool is the number vLLM reports at --max-model-len 32768,
--gpu-memory-utilization 0.92.
What this table does not cover:
- The BF16 original. It is 156.2 GB and does not fit on one card, so Aleph Alpha's FP8 release is the reference row.
- German. Both suites are English. Kolibri is built for German and English, and we did not measure German.
- Context past 32K, and anything but Blackwell. The release is validated to 1M tokens; we did not test the 4-bit experts there. NVFP4 runs natively on Blackwell only.
Comparable with our other models
Accuracy numbers move for reasons that have nothing to do with the model: a shorter token budget, a
different temperature, or whether the model was allowed to reason at all. So every number in this
table, on this card and on our other cards, comes from the one fixed protocol described above, the
same 1,370 items, auto-scored, no LLM judge.
| model | shape | size | overall | knowledge | call | abstain | finished | out/answer |
|---|---|---|---|---|---|---|---|---|
| Laguna-XS-2.1 | 31 B MoE | 19.3 GiB | 81.7 | 83.8 | 68.4 | 73.5 | 98.9% | 1097 |
| Nemotron-3.5-Lightning-30B-A3B | 30 B MoE+Mamba | 19.2 GiB | 87.1 | 87.9 | 85.4 | 70.5 | 97.9% | 1429 |
| Ornith-1.5-35B-A3B | 35 B MoE | 22.6 GiB | 88.7 | 91.7 | 74.4 | 60.0 | 99.3% | 760 |
| Muse-Glimmer-30B | 30 B MoE | 20.4 GiB | 86.6 | 88.8 | 78.6 | 54.5 | 99.7% | 800 |
| Qwen3.8-27B | 27 B dense | 20.7 GiB | 88.8 | 90.4 | 85.5 | 54.5 | 99.7% | 651 |
| Granite-4.2-30B | 30 B dense | 18.1 GB | 85.5 | 86.2 | 85.8 | 60.8 | 98.5% | 1502 |
| Nex-N2.5-mini NVFP4 | 35 B MoE, 3 B active | 23.91 GB | 88.5 | 90.5 | 81.9 | 57.5 | 99.3% | 524 |
| Nex-N2.5-mini mixed | 35 B MoE, 3 B active | 26.04 GB | 88.5 | 90.6 | 80.9 | 58.7 | 99.2% | 504 |
| Nex-N2.5-mini FP8 | 35 B MoE, 3 B active | 38.13 GB | 88.7 | 90.9 | 79.4 | 60.0 | 99.5% | 545 |
| K2-Horizon-MoVA-36B-A4B NVFP4 | 37 B MoE+MoVA, 4 B active | 36.7 GB | 84.1 | 86.5 | 71.8 | 60.4 | 95.6% | 1234 |
| K2-Horizon-MoVA-36B-A4B mixed | 37 B MoE+MoVA, 4 B active | 44.5 GB | 84.9 | 87.3 | 73.5 | 58.9 | 96.3% | 1118 |
| Laguna-S-2.1 | 110 B MoE | 64.0 GiB | 84.3 | 87.1 | 64.6 | 81.0 | 97.3% | 995 |
| Qwen3.8-Flash-Next | 180 B MoE, 6 B active | 183.7 GB | 90.3 | 92.2 | 84.8 | 56.7 | 99.5% | 686 |
| MiMo-V2.6-Flash-RL NVFP4 | 309 B MoE, 15 B active | 182.40 GB | 89.3 | 91.1 | 83.0 | 61.0 | 99.0% | 664 |
| Kolibri-1 FP8 | 78 B MoE, 3.5 B active | 78.83 GB | 85.0 | 86.1 | 86.1 | 48.0 | 97.0% | 1824 |
| Kolibri-1 NVFP4 | 78 B MoE, 3.5 B active | 47.70 GB | 84.9 | 86.2 | 84.9 | 48.0 | 96.4% | 1947 |
| Kolibri-1 mixed (this repo) | 78 B MoE, 3.5 B active | 45.81 GB | 83.9 | 85.0 | 85.5 | 47.5 | 96.6% | 1848 |
overall pools the two suites as 1,370 items, weighted 85.4% knowledge and 14.6% tool calling by
item count. Read it with finished: overall scores an answer that overran the token budget as
wrong, and cannot say whether the model needed the room or failed to stop. A gap under 1.0 in
overall is a tie. Sizes are as each card reports them, which mixes GB and GiB.
What's quantized to what
Kolibri-1 is 78.10 B parameters, 3.46 B active per token, and 96.7% of them are routed experts.
| params | share | |
|---|---|---|
| routed experts, 50 layers × 384 × (gate+up+down) | 75.50 B | 96.7% |
| attention, 50 layers (q, k, v, o) | 1.70 B | 2.2% |
embed_tokens and lm_head, untied, vocab 128000 |
0.66 B | 0.8% |
| shared experts | 0.20 B | 0.3% |
| routers, norms, expert bias | 0.05 B | 0.1% |
| tensors | count | format |
|---|---|---|
| routed experts on all 50 layers | 57,600 modules | NVFP4: E2M1 codes, E4M3 scale per 16, FP32 global, W4A4 |
| attention and shared experts | 350 weights | FP8 E4M3, 128×128 block scales, byte-identical to the release |
| everything else | 403 tensors | BF16, byte-identical to the release |
compressed-tensors, two config groups, each with its own format: nvfp4-pack-quantized for the
experts and float-quantized for attention and shared experts, with dynamic per-group FP8
activations. The release calls its block scale weight_scale_inv and compressed-tensors calls it
weight_scale; the values are the same. vLLM only treats a group as activation-quantized when the
group names one of its formats, so the top-level mixed-precision label alone fails at load with
No compressed-tensors compatible scheme was found.
gate and up of every expert share one global scale, so the fused projection vLLM builds keeps the exact values. All 58,353 tensors of the BF16 source are accounted for, and loading logged no scale warning on any layer. Expert weights are round-to-nearest from the BF16 release, with no calibration.
![]()
primitive ·
more models ·
inference economics for production LLM systems
- Downloads last month
- 37
Model tree for primitive-ai/Kolibri-1-mixed-NVFP4-FP8
Base model
Aleph-Alpha/Kolibri-1-BF16