TandemLLM-Kolibri-1-NVFP4
A mixed NVFP4 / FP8 quantisation of Aleph-Alpha/Kolibri-1, built for the TandemLLM engine on an NVIDIA DGX Spark (GB10). It is 44.9 GB and reads 2.44 GB of weights per decoded token at 1k context.
This is not a drop-in checkpoint for transformers or vLLM. It runs only on the kolibri-experimental branch of TandemLLM. We stopped that work on 2026-10-04: the branch is archived, not maintained, and will not be merged. docs/kolibri.md shows how to run it and what we measured.
Formats
| Part | Format |
|---|---|
| Routed experts | NVFP4: e2m1 codes, e4m3 scales per 16 values, one fp32 scale per tensor |
| q and o of the 40 sliding-window layers | NVFP4 |
| k and v of every layer | FP8 e4m3, 128×128 blocks, fp32 block scales |
| All attention of the 10 full-attention layers (4, 9, …, 49) | FP8 e4m3, 128×128 blocks, fp32 block scales |
| Shared expert | FP8 e4m3, 128×128 blocks, fp32 block scales |
| LM head | e4m3, one fp32 scale per row, fp32 logits |
| Router, norms, embedding | BF16 |
A projection with weight_scale_inv is FP8; a projection with weight_scale_2 is NVFP4. Apply the FP8 block scales in fp32: rounding them to BF16 measurably raises the loss.
How we built it
- Source: the BF16 checkpoint. The FP8 tensors equal the FP8 release's rounding of those BF16 weights.
- The NVFP4 tensors come from GPTQ with clip search. The reference for the error is the FP8 model, because Aleph Alpha trained Kolibri with FP8 quantisation-aware RL.
- Calibration used about 600k tokens of English, code and German (about 30 % German). Experts that the router rarely picks fall back to clip-only rounding.
Quality
Mean NLL difference in nats per token against the FP8 model, pooled per domain (block standard error in brackets). The pass bar was +0.05 with a 0.01 margin.
| Domain | vs FP8 | vs BF16 |
|---|---|---|
| English | +0.035 (0.003) | +0.021 |
| Code | +0.012 (0.005) | +0.000 |
| German | +0.025 (0.003) | +0.022 |
| Chat (34 ChatML sequences) | +0.005 (0.001) | +0.004 |
Five greedy test prompts all end on their own end token. The tokens match the FP8 model's top choice at 92 % to 100 % of positions.
Files
layers/0.safetensors…layers/49.safetensors: one file per decoder layeroutside.safetensors: embedding, final norm, LM headmanifest.json: format, shapes, sha256 and size of every filesummary.json: the quality-check resultsconfig.json,generation_config.json,tokenizer.json,tokenizer_config.json,LICENSE: copied unchanged from the base model
License
Apache-2.0, the same as the base model. Kolibri-1 is by Aleph Alpha; see their tech report and the base model card.
- Downloads last month
- 32
Model tree for 0xBakeer/TandemLLM-Kolibri-1-NVFP4
Base model
Aleph-Alpha/Kolibri-1-BF16