Configuration Parsing Warning:Invalid JSON for config file config.json
NVIDIA-Nemotron-3-Nano-4B-BF16 — NAS-Pruned
Combined depth + width pruned variant of NVIDIA-Nemotron-3-Nano-4B-BF16, found via Neural Architecture
Search targeting ~3.5B total parameters. Rather than manually picking layer count /
hidden size / FFN size, NAS jointly searches this space and evaluates top_k
candidates to pick the best-performing configuration at the target size.
Base model: Hybrid transformer/Mamba-SSM decoder-only LLM, 42 layers, hidden_size=3136, ffn_hidden_size=12544, 40 attention heads, 8 GQA key-value groups (~4.4B total params).
Compression Method
torchrun --nproc_per_node 1 prune_minitron.py \
--hf_model_name_or_path checkpoints/NVIDIA-Nemotron-3-Nano-4B-BF16 \
--calib_dataset_name wikitext \
--calib_num_samples 512 \
--prune_target_params 3.5e9 \
--output_hf_path pruning/NVIDIA-Nemotron-3-Nano-4B-BF16-pruned-NAS/ \
--top_k 5 \
--trust_remote_code
Evaluation
Evaluated 5-shot with lm_eval on WinoGrande, ARC-Easy, and PIQA:
| Task | Score |
|---|---|
| winogrande | 0.6022 |
| arc_easy | 0.6721 |
| piqa | 0.6931 |
For reference, the full BF16 teacher (NVIDIA-Nemotron-3-Nano-4B-BF16) scores 5-shot:
| Task | Score |
|---|---|
| winogrande | 0.6717 |
| arc_easy | 0.8136 |
| piqa | 0.7715 |
Framework
Produced with NVIDIA Model-Optimizer / Megatron-Bridge, as part of the DLI course "The Art of Compressing LLMs: Pruning, Distillation, and Quantization Demystified."
License
Inherits the license terms of the base model, nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16. Check the original
model card for the applicable NVIDIA Open Model License terms before redistribution.
- Downloads last month
- 3