Recipe:
Main model:
Routed Expert Projections - NVFP4 (calibrated, with input scales for A4)
Attention Projections: BF16
Shared Experts: BF16
MTP:
Routed Expert Projections - BF16
Attention Projections: BF16
Shared Experts: BF16
Should fit on 4x RTX 6000. See the slightly smaller -4p67 model for 2x RTX 6000.
KLD: ~0.04
Evals to follow.
Should be used with https://github.com/local-inference-lab/vllm/tree/dev/jovian-judgement