Configuration Parsing Warning:Invalid JSON for config file config.json

NVIDIA-Nemotron-3-Nano-4B-BF16 — NAS-Pruned

Combined depth + width pruned variant of NVIDIA-Nemotron-3-Nano-4B-BF16, found via Neural Architecture Search targeting ~3.5B total parameters. Rather than manually picking layer count / hidden size / FFN size, NAS jointly searches this space and evaluates top_k candidates to pick the best-performing configuration at the target size.

Base model: Hybrid transformer/Mamba-SSM decoder-only LLM, 42 layers, hidden_size=3136, ffn_hidden_size=12544, 40 attention heads, 8 GQA key-value groups (~4.4B total params).

Compression Method

torchrun --nproc_per_node 1 prune_minitron.py \
    --hf_model_name_or_path checkpoints/NVIDIA-Nemotron-3-Nano-4B-BF16 \
    --calib_dataset_name wikitext \
    --calib_num_samples 512 \
    --prune_target_params 3.5e9 \
    --output_hf_path pruning/NVIDIA-Nemotron-3-Nano-4B-BF16-pruned-NAS/ \
    --top_k 5 \
    --trust_remote_code

Evaluation

Evaluated 5-shot with lm_eval on WinoGrande, ARC-Easy, and PIQA:

Task Score
winogrande 0.6022
arc_easy 0.6721
piqa 0.6931

For reference, the full BF16 teacher (NVIDIA-Nemotron-3-Nano-4B-BF16) scores 5-shot:

Task Score
winogrande 0.6717
arc_easy 0.8136
piqa 0.7715

Framework

Produced with NVIDIA Model-Optimizer / Megatron-Bridge, as part of the DLI course "The Art of Compressing LLMs: Pruning, Distillation, and Quantization Demystified."

License

Inherits the license terms of the base model, nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16. Check the original model card for the applicable NVIDIA Open Model License terms before redistribution.

Downloads last month
3
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support