size 47.70 GB 1.65x smaller than the FP8 release accuracy level with the release 1.09x the release throughput at concurrency 32 KV pool 1.40M tokens requires Blackwell primitive.com

Level with Aleph Alpha's FP8 release on accuracy, at 60% of its size.

NVFP4 build of Aleph-Alpha/Kolibri-1-BF16 at 47.70 GB, on one 96 GB card.
All 50 layers of routed experts at 4 bits; attention, shared experts, embeddings, lm_head, routers and norms stay BF16.

Two builds, one measurement. We quantized the experts two ways and ran both against Aleph Alpha's own FP8 release under one protocol. This build is the more accurate on knowledge. The mixed build keeps the release's FP8 attention instead of BF16, and is the faster of the two. On tool calling all three are one band.


Why this quant

  • 🗜️ 47.70 GB against 78.83 GB for the FP8 release. Weights take 44.53 GiB of the card against 73.55 GiB, and the KV pool at 32K context goes from 410,315 tokens to 1,400,826 (3.4×).
  • 🎯 Level with the release. Knowledge 86.2 against 86.1, tool calling 77.5 against 78.5 over five runs each. Both gaps sit inside the run-to-run spread.
  • ⚡ 1,076 tok/s at concurrency 32, against 986 for the release (+9%). Single-stream it is slower: 124.4 tok/s against 159.1.
  • 🐢 BF16 attention costs single-stream speed. The mixed build, which keeps the release's FP8 attention and shared experts, decodes 15% faster at concurrency 1 and 11% faster at 32, and is 1.89 GB smaller. It gives up about a point of knowledge for that.
  • 🧩 Stock Aleph Alpha image, nothing patched. compressed-tensors, nvfp4-pack-quantized.

Serve it

hf download primitive-ai/Kolibri-1-NVFP4 --local-dir ./Kolibri-1-NVFP4

docker run --gpus all --ipc=host -p 8000:8000 -v $PWD:/models \
  ghcr.io/aleph-alpha/aleph-alpha-inference:1.0.0-vllm0.29.0 \
  /models/Kolibri-1-NVFP4 --served-model-name kolibri \
  --kv-cache-dtype fp8 --max-model-len 32768 --gpu-memory-utilization 0.92 \
  --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice

Stock vLLM does not know the Kolibri1ForCausalLM architecture. Aleph Alpha's plugin image adds it along with the kolibri1 parsers, and nothing here is patched on top of it. The image is vLLM 0.29.0 with aleph-alpha-inference 1.0.0 (sha256:9a56ab1691f8bc8eb0fc82f0f534bfdd36da2218fa79b01b83448739fbda403b). Boot takes about 151 s on an idle card, and the NVFP4 experts run on FlashInfer CUTLASS.

Aleph Alpha recommends temperature 1.0, top_p 0.97, top_k 128. Our numbers use the fixed protocol above for every build, so they are comparable with each other, not with Aleph Alpha's own. The measurement servers ran without the two parsers, because the tool-calling suite does not pass tools=. Thinking is on by default; send enable_thinking: false to turn it off.


Measured

One RTX PRO 6000 Blackwell, 96 GB, one card. The 1,170-item knowledge suite and the 200-item tool-calling suite, temperature 0.6 / top_p 0.95 / top_k 20, thinking on, a 16,384-token budget, concurrency 32, auto-scored with no LLM judge. FP8 KV cache for every row, which is how Aleph Alpha evaluates the release. Throughput is 8K in / 512 out, prefix-cache free, two seeds per cell. Every row ran on the same box on the same day.

build size knowledge tool-calling call abstain finished tok/s @1 tok/s @32 KV pool
Kolibri-1, FP8 release 78.83 GB 86.1 78.5 86.1 48.0 97.0% 159.1 986 410,315
mixed 45.81 GB 85.0 77.9 85.5 47.5 96.6% 142.6 1195 1,461,977
this repo, NVFP4 47.70 GB 86.2 77.5 84.9 48.0 96.4% 124.4 1076 1,400,826

Tool-calling is the mean of five runs per build, with a within-build standard deviation of 0.6 to 1.2, so that column is one band. Knowledge is two runs per build (86.0/86.2, 84.8/85.1, 85.9/86.5 strict, in table order). On knowledge the release and NVFP4 are 0.1 apart; the mixed build is 1.1 below the release and 1.2 below NVFP4, on both of its runs. KV pool is the number vLLM reports at --max-model-len 32768, --gpu-memory-utilization 0.92.

What this table does not cover:

  • The BF16 original. It is 156.2 GB and does not fit on one card, so Aleph Alpha's FP8 release is the reference row.
  • German. Both suites are English. Kolibri is built for German and English, and we did not measure German.
  • Context past 32K, and anything but Blackwell. The release is validated to 1M tokens; we did not test the 4-bit experts there. NVFP4 runs natively on Blackwell only.

Comparable with our other models

Accuracy numbers move for reasons that have nothing to do with the model: a shorter token budget, a different temperature, or whether the model was allowed to reason at all. So every number in this table, on this card and on our other cards, comes from the one fixed protocol described above, the same 1,370 items, auto-scored, no LLM judge.

model shape size overall knowledge call abstain finished out/answer
Laguna-XS-2.1 31 B MoE 19.3 GiB 81.7 83.8 68.4 73.5 98.9% 1097
Nemotron-3.5-Lightning-30B-A3B 30 B MoE+Mamba 19.2 GiB 87.1 87.9 85.4 70.5 97.9% 1429
Ornith-1.5-35B-A3B 35 B MoE 22.6 GiB 88.7 91.7 74.4 60.0 99.3% 760
Muse-Glimmer-30B 30 B MoE 20.4 GiB 86.6 88.8 78.6 54.5 99.7% 800
Qwen3.8-27B 27 B dense 20.7 GiB 88.8 90.4 85.5 54.5 99.7% 651
Granite-4.2-30B 30 B dense 18.1 GB 85.5 86.2 85.8 60.8 98.5% 1502
Nex-N2.5-mini NVFP4 35 B MoE, 3 B active 23.91 GB 88.5 90.5 81.9 57.5 99.3% 524
Nex-N2.5-mini mixed 35 B MoE, 3 B active 26.04 GB 88.5 90.6 80.9 58.7 99.2% 504
Nex-N2.5-mini FP8 35 B MoE, 3 B active 38.13 GB 88.7 90.9 79.4 60.0 99.5% 545
K2-Horizon-MoVA-36B-A4B NVFP4 37 B MoE+MoVA, 4 B active 36.7 GB 84.1 86.5 71.8 60.4 95.6% 1234
K2-Horizon-MoVA-36B-A4B mixed 37 B MoE+MoVA, 4 B active 44.5 GB 84.9 87.3 73.5 58.9 96.3% 1118
Laguna-S-2.1 110 B MoE 64.0 GiB 84.3 87.1 64.6 81.0 97.3% 995
Qwen3.8-Flash-Next 180 B MoE, 6 B active 183.7 GB 90.3 92.2 84.8 56.7 99.5% 686
MiMo-V2.6-Flash-RL NVFP4 309 B MoE, 15 B active 182.40 GB 89.3 91.1 83.0 61.0 99.0% 664
Kolibri-1 FP8 78 B MoE, 3.5 B active 78.83 GB 85.0 86.1 86.1 48.0 97.0% 1824
Kolibri-1 mixed 78 B MoE, 3.5 B active 45.81 GB 83.9 85.0 85.5 47.5 96.6% 1848
Kolibri-1 NVFP4 (this repo) 78 B MoE, 3.5 B active 47.70 GB 84.9 86.2 84.9 48.0 96.4% 1947

overall pools the two suites as 1,370 items, weighted 85.4% knowledge and 14.6% tool calling by item count. Read it with finished: overall scores an answer that overran the token budget as wrong, and cannot say whether the model needed the room or failed to stop. A gap under 1.0 in overall is a tie. Sizes are as each card reports them, which mixes GB and GiB.


What's quantized to what

Kolibri-1 is 78.10 B parameters, 3.46 B active per token, and 96.7% of them are routed experts.

params share
routed experts, 50 layers × 384 × (gate+up+down) 75.50 B 96.7%
attention, 50 layers (q, k, v, o) 1.70 B 2.2%
embed_tokens and lm_head, untied, vocab 128000 0.66 B 0.8%
shared experts 0.20 B 0.3%
routers, norms, expert bias 0.05 B 0.1%
tensors count format
routed experts on all 50 layers 57,600 modules NVFP4: E2M1 codes, E4M3 scale per 16, FP32 global, W4A4
everything else 753 tensors BF16, byte-identical to the BF16 source

compressed-tensors, format nvfp4-pack-quantized, one config group, W4A4 with group-16 weight scales and a per-module global scale. The source already ships experts as per-expert modules, so there is nothing to unfold. gate and up of every expert share one weight_global_scale; vLLM fuses those two halves and keeps one scale, warning and taking the maximum when they disagree, and loading logged no such warning on any layer. All 58,353 tensors of the BF16 source are accounted for. Expert weights are round-to-nearest, with no calibration.



Primitive
primitive · more models · inference economics for production LLM systems

Downloads last month
105
Safetensors
Model size
78B params
Tensor type
U8
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for primitive-ai/Kolibri-1-NVFP4

Quantized
(14)
this model