TandemLLM-Kolibri-1-NVFP4

A mixed NVFP4 / FP8 quantisation of Aleph-Alpha/Kolibri-1, built for the TandemLLM engine on an NVIDIA DGX Spark (GB10). It is 44.9 GB and reads 2.44 GB of weights per decoded token at 1k context.

This is not a drop-in checkpoint for transformers or vLLM. It runs only on the kolibri-experimental branch of TandemLLM. We stopped that work on 2026-10-04: the branch is archived, not maintained, and will not be merged. docs/kolibri.md shows how to run it and what we measured.

Formats

Part Format
Routed experts NVFP4: e2m1 codes, e4m3 scales per 16 values, one fp32 scale per tensor
q and o of the 40 sliding-window layers NVFP4
k and v of every layer FP8 e4m3, 128×128 blocks, fp32 block scales
All attention of the 10 full-attention layers (4, 9, …, 49) FP8 e4m3, 128×128 blocks, fp32 block scales
Shared expert FP8 e4m3, 128×128 blocks, fp32 block scales
LM head e4m3, one fp32 scale per row, fp32 logits
Router, norms, embedding BF16

A projection with weight_scale_inv is FP8; a projection with weight_scale_2 is NVFP4. Apply the FP8 block scales in fp32: rounding them to BF16 measurably raises the loss.

How we built it

  • Source: the BF16 checkpoint. The FP8 tensors equal the FP8 release's rounding of those BF16 weights.
  • The NVFP4 tensors come from GPTQ with clip search. The reference for the error is the FP8 model, because Aleph Alpha trained Kolibri with FP8 quantisation-aware RL.
  • Calibration used about 600k tokens of English, code and German (about 30 % German). Experts that the router rarely picks fall back to clip-only rounding.

Quality

Mean NLL difference in nats per token against the FP8 model, pooled per domain (block standard error in brackets). The pass bar was +0.05 with a 0.01 margin.

Domain vs FP8 vs BF16
English +0.035 (0.003) +0.021
Code +0.012 (0.005) +0.000
German +0.025 (0.003) +0.022
Chat (34 ChatML sequences) +0.005 (0.001) +0.004

Five greedy test prompts all end on their own end token. The tokens match the FP8 model's top choice at 92 % to 100 % of positions.

Files

  • layers/0.safetensors … layers/49.safetensors: one file per decoder layer
  • outside.safetensors: embedding, final norm, LM head
  • manifest.json: format, shapes, sha256 and size of every file
  • summary.json: the quality-check results
  • config.json, generation_config.json, tokenizer.json, tokenizer_config.json, LICENSE: copied unchanged from the base model

License

Apache-2.0, the same as the base model. Kolibri-1 is by Aleph Alpha; see their tech report and the base model card.

Downloads last month
32
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0xBakeer/TandemLLM-Kolibri-1-NVFP4

Quantized
(17)
this model