Qwen3-30B-A3B INT8 for AWS Inferentia2 (Archived)
โ ๏ธ Deployment Status: Archived โ too large for inf2.xlarge; use smaller models.
INT8-quantized Qwen3-30B-A3B (MoE) compiled for AWS Inferentia2.
Model Details
- Base model: Qwen/Qwen3-30B-A3B
- Architecture: Mixture of Experts (MoE), 30B total / 3B active per token
- Quantization: INT8 per-channel symmetric (axis=2 for expert layers)
- Target hardware: AWS Inferentia2
MoE INT8 Quantization Notes
Standard per_tensor_symmetric quantization breaks on MoE expert layers.
Expert weights require per_channel_symmetric with axis=2, and expert
scale tensors must be stacked manually:
# For MoE expert layers:
# scale shape: [num_experts, out_dim, 1]
expert_scales = torch.stack([expert_scales[i] for i in range(num_experts)])
Why Archived
30B MoE checkpoint exceeds CPU RAM capacity of inf2.xlarge during weight loading
(DMA-pinned, cannot be swapped). Requires inf2.8xlarge or larger.
Related Models
- aqidd/qwen3-8b-int8-inf2 โ production-ready on inf2.xlarge
- aqidd/qwen3-14b-int8-inf2 โ production-ready on inf2.xlarge
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support