NVIDIA Nemotron 3.5 Lightning 30B-A3B — APEX GGUF

Imatrix-guided, measured-allocation APEX quantizations of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B — a Nemotron-H hybrid: 52 layers of which only 6 are full attention, the rest split between Mamba2 SSM and MoE (128 routed experts + shared expert, ~3B active of 31.6B).

Three things distinguish these from the other GGUFs of this model:

  • The MTP head is included. The checkpoint ships a multi-token-prediction head; most quants drop it, which makes speculative decoding impossible. These keep it (53 blocks, not 52).
  • Built with an importance matrix, generated on this architecture specifically. The others are stock conversions.
  • Per-tensor bit allocation is measured, not assumed — each tensor was probed for how much output KL it actually costs at each candidate width, and the budget spent accordingly.

Files

tier size bpw wikitext PPL vs bf16 for status
mini 13.66 GiB 3.56 7.843 +11.7% 16 GB card with the full 128k context available
compact 16.91 GiB 4.42 7.252 +3.3% the value pick — 5 GB smaller for 2.8% uploading
i-quality 21.53 GiB 5.62 7.053 +0.46% best quality; 24 GB card or unified memory uploading
bf16 reference 61.32 GiB 16.0 7.021 (not hosted — measured as the baseline)

mini is up now; compact and i-quality are still being uploaded. All three are built, measured and gated — the numbers above are from the finished files — but only what the file listing shows is downloadable yet.

nemotron-lightning.imatrix (56 MB) is included so you can build your own tiers.

Reproducing the perplexity numbers

llama-perplexity -m <tier>.gguf -f wiki.test.raw

wiki.test.raw is the unmodified WikiText-2 raw test split — 1,292,013 bytes, 241,211 words, the file llama.cpp's own perplexity documentation uses, so these numbers are directly comparable to anyone else's. From wikitext-2-raw-v1, which shares its test split with WikiText-103. Identical settings across all four rows above; default context.

The calibration corpus is not WikiText. The imatrix was built on general prose and scientific text, deliberately, so the reported perplexity is measured on data the quantization never saw. Calibrating on WikiText train and then scoring on WikiText test flatters the result — the splits come from the same distribution — and it is an easy mistake to make, since the obvious calibration file to reach for is often exactly that.

Why mini fits a 16 GB card when a 30B usually doesn't

This model spends 6 KiB per token of KV cache — only 6 of 52 layers are attention, and the Mamba state is constant-size regardless of context length. So the full 128k context costs 0.75 GiB:

13.66 GiB weights + 0.75 GiB KV @ 128k = 14.41 GiB

For comparison, a conventional 35B-class MoE at ~82 KiB/token would need 10.5 GiB for that same context — the cache alone would not fit the card, let alone the weights. That pairing is the reason this model is interesting at this size point.

Quality gate

Every tier published here passed all three checks. Nothing is uploaded that did not.

tier PPL ratio (bar: ≤1.50) coherent generation chained tool-calling
i-quality 1.00 pass 5/5
compact 1.03 pass 5/5
mini 1.12 pass 5/5

Tool-calling is a two-turn dependent chain with a distractor tool that must not be called — a single trivial call is too easy to discriminate between tiers.

Usage

llama-server -m NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-mini.gguf \
  --ctx-size 131072 -fa on --jinja

The chat template is embedded, so --jinja is enough. Always pass --ctx-size — it otherwise defaults to the model's trained context.

To use the MTP head for speculative decoding:

llama-server -m ...-APEX-i-quality.gguf --ctx-size 131072 -fa on --jinja \
  --spec-type draft-mtp --spec-draft-n-max 2

Draft depth is model-specific and does not transfer between models — measure it at your own sampling settings and context length rather than copying a number from elsewhere.

This model emits explicit reasoning traces before its answer. Budget output tokens accordingly; a short --n-predict will truncate mid-thought.

Needs a llama.cpp new enough to load the MTP head — if you see a tensor-count mismatch on load, update.

Notes on the architecture

The row dimensions are 2688 (hidden) and 1856 (expert-down), neither divisible by 256. Every k-quant and IQ type requires a 256-wide superblock, so on this model they are all illegal and llama-quantize silently substitutes other types. A stock Q4_K_M of this model measures 6.21 bpw against a nominal 4.85, and contains ~1% actual Q4_K. That is why these tiers use block-32 and block-64 types throughout, chosen deliberately rather than arrived at by fallback.

Attribution

Calibration: a general prose/scientific corpus (no code), matching the other APEX quants in this collection. Unofficial community quantization; not affiliated with or endorsed by NVIDIA.

Downloads last month
413
GGUF
Model size
33B params
Architecture
nemotron_h_moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF

Quantized
(68)
this model