NVIDIA Nemotron Labs 3 Elastic 23B A2.8B GGUF

A smaller, hacker-friendly GGUF build of NVIDIA Nemotron Elastic.

Built for llama.cpp, LM Studio, and other GGUF-compatible runtimes.

Files

File Notes Rough VRAM target
nemotron-elastic-12b-Q4_K_S.gguf Smaller 4-bit quant ~10GB VRAM
nemotron-elastic-12b-Q4_K_M.gguf Better 4-bit quant ~10GB+ VRAM

Which one should I use?

Use Q4_K_S if you want the easier/smaller 4-bit file.

Use Q4_K_M if you want the better-quality 4-bit file and have a little more room.

Both files are intended for roughly 10GB VRAM class hardware, depending on context size, KV cache settings, and GPU offload.

LM Studio

Open LM Studio and search for:

HackerTwins/NVIDIA-Nemotron-Labs-3-Elastic-23B-A2.8B-GGUF
Downloads last month
213
GGUF
Model size
12B params
Architecture
nemotron_h_moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HackerTwins/NVIDIA-Nemotron-Labs-3-Elastic-12B-A2B-GGUF

Collection including HackerTwins/NVIDIA-Nemotron-Labs-3-Elastic-12B-A2B-GGUF