Quantization Details

This model was quantized to 8-bit floating-point (FP8) using the llmcompressor library.

To prioritize fast processing and maintain high accuracy without requiring a calibration dataset, we utilized a zero-shot Round-to-Nearest (RTN) approach with block-wise weight quantization.

Quantization Configuration:

  • Algorithm: Round-to-Nearest (RTN)
  • Weight Scheme: FP8 with a Block Size of 128 (FP8_BLOCK)
  • Activation Scheme: Dynamic FP8 (quantized dynamically during inference)
  • Calibration Data: None (Zero-shot)
  • Preserved Layers: The lm_head and mlp.gate layers were explicitly ignored and kept in original precision to protect the structural integrity and routing accuracy of the model.

Usage: This FP8-quantized model is natively supported by vLLM and transformers. It is highly recommended to run this model in environments optimized for FP8 computation (e.g., NVIDIA Hopper or Ada Lovelace architectures) to achieve the best performance memory-bandwidth reductions.

Downloads last month
6
Safetensors
Model size
32B params
Tensor type
F32
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SaltyYean/sarvam-30b-FP8-BLOCK

Quantized
(30)
this model