Instructions to use 0xSero/Nemotron-3-Super-64B-W4A16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use 0xSero/Nemotron-3-Super-64B-W4A16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="0xSero/Nemotron-3-Super-64B-W4A16", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("0xSero/Nemotron-3-Super-64B-W4A16", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("0xSero/Nemotron-3-Super-64B-W4A16", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use 0xSero/Nemotron-3-Super-64B-W4A16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "0xSero/Nemotron-3-Super-64B-W4A16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "0xSero/Nemotron-3-Super-64B-W4A16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/0xSero/Nemotron-3-Super-64B-W4A16
- SGLang
How to use 0xSero/Nemotron-3-Super-64B-W4A16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "0xSero/Nemotron-3-Super-64B-W4A16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "0xSero/Nemotron-3-Super-64B-W4A16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "0xSero/Nemotron-3-Super-64B-W4A16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "0xSero/Nemotron-3-Super-64B-W4A16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use 0xSero/Nemotron-3-Super-64B-W4A16 with Docker Model Runner:
docker model run hf.co/0xSero/Nemotron-3-Super-64B-W4A16
Support this work → · X · GitHub · REAP paper · Cerebras REAP
Nemotron-3-Super-64B-W4A16
W4A16 quantization of nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16.
At a glance
| Base model | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 |
| Format | W4A16 |
| Total params | 64B |
| Active / token | 12B |
| Experts / layer | 256 |
| Layers | — |
| Hidden size | 4096 |
| Context | 262,144 |
| On-disk size | 42 GB |
Which variant should I pick?
| Variant | Format | Link |
|---|---|---|
Nemotron-3-Super-64B |
BF16 | link |
Nemotron-3-Super-64B-W4A16 (this) |
W4A16 | link |
Nemotron-3-Super-92B |
BF16 | link |
Nemotron-3-Super-92B-W4A16 |
W4A16 | link |
Draft AutoRound quantization of a Nemotron Super checkpoint.
Base model: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
Draft status
This is a draft research release. It is published for inspection, reproducibility, and early runtime validation. It should not be treated as a final benchmarked production checkpoint.
How this was produced
We quantized the checkpoint with Intel AutoRound using the W4A16 scheme on the remote 8x RTX 3090 host. This lane is optimized for overnight completion and resumability rather than final accuracy tuning.
Settings used
- source checkpoint:
/mnt/llm_models/nemotron-super-compressions/nemotron_super_merged_long50_short15120_v2/reap_50pct - source type:
REAP 50% pruned checkpoint - quantizer:
intel/auto-round 0.10.2 - scheme:
W4A16 - format:
auto_round - calibration dataset:
NeelNanda/pile-10k - device_map:
auto - nsamples:
128 - iters:
50 - seqlen:
1024 - batch_size:
2 - nblocks:
1 - low_gpu_mem_usage:
True - output dir:
/home/ser/nemotron-super/autoround_w4a16/reap_50pct
Notes
- upstream provenance is preserved through the base model link above
- this repo is intentionally marked draft while quantization/runtime validation is still in progress
- donation link added per maintainer request
License & citation
License inherited from the base model.
@misc{lasby2025reap,
title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}
Sponsors
Made possible by NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle.
- Downloads last month
- 24