GLM-5.3-Flash-NVFP4 / README.md
lukealonso's picture
Update README.md
378ca54 verified
|
Raw
History Blame Contribute Delete
473 Bytes

Recipe:

  Main model:
  
    Routed Expert Projections - NVFP4 (calibrated, with input scales for A4)
    Attention Projections: BF16
    Shared Experts: BF16
  
  MTP:
    Routed Expert Projections - BF16
    Attention Projections: BF16
    Shared Experts: BF16

Should fit on 4x RTX 6000. See the slightly smaller -4p67 model for 2x RTX 6000.

KLD: ~0.04

Evals to follow.

Should be used with https://github.com/local-inference-lab/vllm/tree/dev/jovian-judgement