Qwen3-0.6B-4bit-qqq-gptq-obq

Qwen3-0.6B, weights quantized to 4-bit (group_size 128, symmetric), 8-bit embedding, ~5.22 bits/weight.

Method: QQQ smoothing + GPTQ within-matrix (act-order) + OBQ across-matrix (GGN + KL-teacher correction). Calibration: gitarist/calibration-generic.

wikitext-2 PPL 23.01, mean KL to fp16 0.153 (fp16 ref PPL 20.96).

With per-token dynamic int8 activations (W4A8): PPL 23.26, KL 0.164.

Weights are provided dequantized in fp16; weight quantization is baked in. The W4A8 numbers require per-token int8 activation quantization applied at inference.

Downloads last month
12
Safetensors
Model size
0.6B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gitarist/Qwen3-0.6B-4bit-qqq-gptq-obq

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1344)
this model