Instructions to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF # Run inference directly in the terminal: ./llama-cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Use Docker
docker model run hf.co/Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
- LM Studio
- Jan
- vLLM
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
- Ollama
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with Ollama:
ollama run hf.co/Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
- Unsloth Studio
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF to start chatting
- Pi
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with Docker Model Runner:
docker model run hf.co/Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
- Lemonade
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Run and chat with the model
lemonade run user.NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF
Run Hermes
hermes
- Atomic Chat
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF# Run inference directly in the terminal:
llama cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUFUse pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF# Run inference directly in the terminal:
./llama-cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUFBuild from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF# Run inference directly in the terminal:
./build/bin/llama-cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUFUse Docker
docker model run hf.co/Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUFNVIDIA Nemotron 3.5 Lightning 30B-A3B — APEX GGUF
Imatrix-guided, measured-allocation APEX quantizations of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B — a Nemotron-H hybrid: 52 layers of which only 6 are full attention, the rest split between Mamba2 SSM and MoE (128 routed experts + shared expert, ~3B active of 31.6B).
Three things distinguish these from the other GGUFs of this model:
- The MTP head is included. The checkpoint ships a multi-token-prediction head; most quants drop it, which makes speculative decoding impossible. These keep it (53 blocks, not 52).
- Built with an importance matrix, generated on this architecture specifically. The others are stock conversions.
- Per-tensor bit allocation is measured, not assumed — each tensor was probed for how much output KL it actually costs at each candidate width, and the budget spent accordingly.
Files
| tier | size | bpw | wikitext PPL | vs bf16 | for | status |
|---|---|---|---|---|---|---|
| mini | 13.66 GiB | 3.56 | 7.843 | +11.7% | 16 GB card with the full 128k context | available |
| compact | 16.91 GiB | 4.42 | 7.252 | +3.3% | the value pick — 5 GB smaller for 2.8% | uploading |
| i-quality | 21.53 GiB | 5.62 | 7.053 | +0.46% | best quality; 24 GB card or unified memory | uploading |
| bf16 reference | 61.32 GiB | 16.0 | 7.021 | — | (not hosted — measured as the baseline) | — |
mini is up now; compact and i-quality are still being uploaded. All three are built, measured and gated — the numbers above are from the finished files — but only what the file listing shows is downloadable yet.
nemotron-lightning.imatrix (56 MB) is included so you can build your own tiers.
Reproducing the perplexity numbers
llama-perplexity -m <tier>.gguf -f wiki.test.raw
wiki.test.raw is the unmodified WikiText-2 raw test split — 1,292,013 bytes, 241,211
words, the file llama.cpp's own perplexity documentation uses, so these numbers are directly
comparable to anyone else's. From
wikitext-2-raw-v1, which shares its
test split with WikiText-103. Identical settings across all four rows above; default context.
The calibration corpus is not WikiText. The imatrix was built on general prose and scientific text, deliberately, so the reported perplexity is measured on data the quantization never saw. Calibrating on WikiText train and then scoring on WikiText test flatters the result — the splits come from the same distribution — and it is an easy mistake to make, since the obvious calibration file to reach for is often exactly that.
Why mini fits a 16 GB card when a 30B usually doesn't
This model spends 6 KiB per token of KV cache — only 6 of 52 layers are attention, and the Mamba state is constant-size regardless of context length. So the full 128k context costs 0.75 GiB:
13.66 GiB weights + 0.75 GiB KV @ 128k = 14.41 GiB
For comparison, a conventional 35B-class MoE at ~82 KiB/token would need 10.5 GiB for that same context — the cache alone would not fit the card, let alone the weights. That pairing is the reason this model is interesting at this size point.
Quality gate
Every tier published here passed all three checks. Nothing is uploaded that did not.
| tier | PPL ratio (bar: ≤1.50) | coherent generation | chained tool-calling |
|---|---|---|---|
| i-quality | 1.00 | pass | 5/5 |
| compact | 1.03 | pass | 5/5 |
| mini | 1.12 | pass | 5/5 |
Tool-calling is a two-turn dependent chain with a distractor tool that must not be called — a single trivial call is too easy to discriminate between tiers.
Usage
llama-server -m NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-mini.gguf \
--ctx-size 131072 -fa on --jinja
The chat template is embedded, so --jinja is enough. Always pass --ctx-size — it
otherwise defaults to the model's trained context.
To use the MTP head for speculative decoding:
llama-server -m ...-APEX-i-quality.gguf --ctx-size 131072 -fa on --jinja \
--spec-type draft-mtp --spec-draft-n-max 2
Draft depth is model-specific and does not transfer between models — measure it at your own sampling settings and context length rather than copying a number from elsewhere.
This model emits explicit reasoning traces before its answer. Budget output tokens
accordingly; a short --n-predict will truncate mid-thought.
Needs a llama.cpp new enough to load the MTP head — if you see a tensor-count mismatch on load, update.
Notes on the architecture
The row dimensions are 2688 (hidden) and 1856 (expert-down), neither divisible by 256. Every
k-quant and IQ type requires a 256-wide superblock, so on this model they are all illegal and
llama-quantize silently substitutes other types. A stock Q4_K_M of this model measures
6.21 bpw against a nominal 4.85, and contains ~1% actual Q4_K. That is why these tiers use
block-32 and block-64 types throughout, chosen deliberately rather than arrived at by fallback.
Attribution
- Base model: NVIDIA — NVIDIA-Nemotron-3.5-Lightning-30B-A3B
- APEX recipe & toolkit: LocalAI — localai-org/apex-quant
- Quantization engine: llama.cpp (ggml-org)
Calibration: a general prose/scientific corpus (no code), matching the other APEX quants in this collection. Unofficial community quantization; not affiliated with or endorsed by NVIDIA.
- Downloads last month
- 413
We're not able to determine the quantization variants.
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF# Run inference directly in the terminal: llama cli -hf Myric/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-APEX-GGUF