Qwen3-0.6B-3bit-awq-obq

Qwen3-0.6B quantized to 3-bit (group_size 256, symmetric), 8-bit embedding, ~4.40 bits/weight.

Method: AWQ smoothing + OBQ (GGN + KL-teacher correction). Calibration: gitarist/calibration-generic.

wikitext-2 PPL 45.98, mean KL to fp16 0.965 (fp16 ref PPL 20.96).

Weights are provided dequantized in fp16 for direct loading; quantization is baked in — evaluate as-is.

Downloads last month
11
Safetensors
Model size
0.6B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gitarist/Qwen3-0.6B-3bit-awq-obq

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1340)
this model