Nemotron 3.5 Lightning 30B A3B 5GB

The NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 model, quantized to Q4_K_M and split into 5 GB GGUF files.

5GB split

The model has been split using llama-gguf-split (b10237) as follows:

/llama-gguf-split --split-max-size 5G NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M
Downloads last month
103
GGUF
Model size
33B params
Architecture
nemotron_h_moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-snaps/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-5GB

Quantized
(68)
this model