--- license: apache-2.0 language: - id - en tags: - aws-inferentia2 - neuron - int8 - qwen3 - moe - archived --- # Qwen3-30B-A3B INT8 for AWS Inferentia2 (Archived) > ⚠️ **Deployment Status: Archived** — too large for inf2.xlarge; use smaller models. INT8-quantized Qwen3-30B-A3B (MoE) compiled for AWS Inferentia2. ## Model Details - **Base model**: [Qwen/Qwen3-30B-A3B](https://huggingface.co/Qwen/Qwen3-30B-A3B) - **Architecture**: Mixture of Experts (MoE), 30B total / 3B active per token - **Quantization**: INT8 per-channel symmetric (axis=2 for expert layers) - **Target hardware**: AWS Inferentia2 ## MoE INT8 Quantization Notes Standard `per_tensor_symmetric` quantization breaks on MoE expert layers. Expert weights require `per_channel_symmetric` with `axis=2`, and expert scale tensors must be stacked manually: ```python # For MoE expert layers: # scale shape: [num_experts, out_dim, 1] expert_scales = torch.stack([expert_scales[i] for i in range(num_experts)]) ``` ## Why Archived 30B MoE checkpoint exceeds CPU RAM capacity of inf2.xlarge during weight loading (DMA-pinned, cannot be swapped). Requires `inf2.8xlarge` or larger. ## Related Models - [aqidd/qwen3-8b-int8-inf2](https://huggingface.co/aqidd/qwen3-8b-int8-inf2) — production-ready on inf2.xlarge - [aqidd/qwen3-14b-int8-inf2](https://huggingface.co/aqidd/qwen3-14b-int8-inf2) — production-ready on inf2.xlarge