Text Generation
Transformers
GGUF
PyTorch
nvidia
nemotron-3.5
imatrix
conversational

Could you upload imatrix_unsloth.gguf for Lightning? (+ two questions about the MTP layer)

#2
by pirola - opened

Hi β€” thanks for getting Lightning quants out so fast.

  1. The imatrix file seems to be missing from the repo. Your own quant metadata references it:

quantize.imatrix.file = NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF/imatrix_unsloth.gguf
quantize.imatrix.dataset = unsloth_calibration_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.txt
quantize.imatrix.chunks_count = 80
quantize.imatrix.entries_count = 185

but there's no imatrix_unsloth.gguf in the file listing (you shipped one for Nemotron-3-Nano-30B-A3B). Was it just missed in the upload? Having it would let people reproduce or re-mix quants without recomputing calibration.

  1. Does the imatrix cover blk.52 (the MTP/nextn layer)? With entries_count = 185 I suspect it doesn't. Using a third-party imatrix, llama-quantize bails with:

Missing importance matrix for tensor blk.52.ffn_down_exps.weight in a very low-bit quantization

I notice your files put blk.52.ffn_{up,down}_exps at Q5_0 β€” a block-32 type that needs no importance data β€” while the trunk experts get the imatrix-guided types. Was that a deliberate workaround for the same issue, or does your imatrix actually include blk.52?

  1. Is there any way to get importance data for the MTP head? As far as I can tell llama-imatrix never exercises the nextn path (it's not part of the normal forward pass), so those tensors can't get statistics by construction. Is there a flag or procedure that works, or is high-precision-for-blk.52 simply the right answer?

  2. Does MTP speculative decoding actually work for this model? i.e. --spec-type draft-mtp β€” have you measured a speedup, and does it need anything beyond keeping blk.52 in the file? You kept the whole MTP head at high precision, which suggests it's meant to be usable.

Context, in case it's useful: I'm quantizing this model with zero-padded expert tensors (1856β†’2048, 2688β†’2816) so block-256 quants actually apply β€” the Nemotron-3 family's expert dims aren't 256-divisible, which is why (for example) your UD-IQ3_XXS ends up with the 46 trunk expert tensors as IQ4_NL rather than IQ3_XXS, and lands at 19.8 GB. Padded, the same recipe gives genuine IQ3_XXS experts at ~13.7 GB. Happy to share details if that's interesting to you.

I did it successfully with nemotron 3 nano: pirola/Nemotron-3-Nano-30B-A3B-pirola-IQ2_XXS-XS-GGUF and pirola/Nemotron-3-Nano-30B-A3B-pirola-IQ3_XXS-GGUF
i did the pirola/Nemotron-3.5-Lightning-30B-A3B-pirola-IQ3_XXS-GGUF but would rather use your imatrix

pirola changed discussion title from imatrix please? to Could you upload imatrix_unsloth.gguf for Lightning? (+ two questions about the MTP layer)

@danielhanchen could you please help me here?

Sign up or log in to comment