Qwen3-30B-A3B INT8 for AWS Inferentia2 (Archived)

โš ๏ธ Deployment Status: Archived โ€” too large for inf2.xlarge; use smaller models.

INT8-quantized Qwen3-30B-A3B (MoE) compiled for AWS Inferentia2.

Model Details

  • Base model: Qwen/Qwen3-30B-A3B
  • Architecture: Mixture of Experts (MoE), 30B total / 3B active per token
  • Quantization: INT8 per-channel symmetric (axis=2 for expert layers)
  • Target hardware: AWS Inferentia2

MoE INT8 Quantization Notes

Standard per_tensor_symmetric quantization breaks on MoE expert layers. Expert weights require per_channel_symmetric with axis=2, and expert scale tensors must be stacked manually:

# For MoE expert layers:
# scale shape: [num_experts, out_dim, 1]
expert_scales = torch.stack([expert_scales[i] for i in range(num_experts)])

Why Archived

30B MoE checkpoint exceeds CPU RAM capacity of inf2.xlarge during weight loading (DMA-pinned, cannot be swapped). Requires inf2.8xlarge or larger.

Related Models

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support