variant: id: onnx_int8 format: onnx precision: int8 # INT8 quantized — auto-cast by NPU execution provider at runtime method: onnx_export # base model exported to ONNX; quantization applied by runtime quantization: datatype: int8 scope: - weights - activations granularity: per-tensor # typical for runtime-cast INT8 calibration: ptq # post-training quantization applied by the NPU EP / MWMX toolchain toolchain: onnx