Quantization Details
This model was quantized to 8-bit floating-point (FP8) using the llmcompressor library.
To prioritize fast processing and maintain high accuracy without requiring a calibration dataset, we utilized a zero-shot Round-to-Nearest (RTN) approach with block-wise weight quantization.
Quantization Configuration:
- Algorithm: Round-to-Nearest (RTN)
- Weight Scheme: FP8 with a Block Size of 128 (
FP8_BLOCK) - Activation Scheme: Dynamic FP8 (quantized dynamically during inference)
- Calibration Data: None (Zero-shot)
- Preserved Layers: The
lm_headandmlp.gatelayers were explicitly ignored and kept in original precision to protect the structural integrity and routing accuracy of the model.
Usage:
This FP8-quantized model is natively supported by vLLM and transformers. It is highly recommended to run this model in environments optimized for FP8 computation (e.g., NVIDIA Hopper or Ada Lovelace architectures) to achieve the best performance memory-bandwidth reductions.
- Downloads last month
- 6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for SaltyYean/sarvam-30b-FP8-BLOCK
Base model
sarvamai/sarvam-30b