Qwen3-1.7B — GPTQ Int4 (group_size 256, symmetric)

A 4-bit GPTQ quantization of Qwen/Qwen3-1.7B. The model keeps the original architecture and tokenizer; only the linear weights in the transformer blocks are quantized to 4-bit.

Quantization format

This checkpoint uses the standard packed GPTQ layout (qweight / qzeros / scales / g_idx per quantized linear).

field value
method gptq
bits 4
group_size 256
symmetric (sym) true
activation reorder (desc_act) false
checkpoint format gptq_v2
recommended backend exllama_v2
Marlin compatible no (Marlin requires group_size 128)

Weights are stored in the gptq_v2 convention, i.e. dequantization is w = scale * (q - zero). Most current loaders (GPTQModel, recent AutoGPTQ, vLLM, Transformers) read the checkpoint_format field in quantize_config.json and select the convention automatically. Because group_size = 256, the Marlin kernel is not eligible; use the exllama_v2 / Triton / CUDA backends instead.

How to run

Transformers + GPTQModel

from transformers import AutoTokenizer
from gptqmodel import GPTQModel

repo = "gitarist/Qwen3-1.7B-GPTQ-Int4-g256-sym"
tok = AutoTokenizer.from_pretrained(repo)
model = GPTQModel.load(repo)

msgs = [{"role": "user", "content": "Explain what 4-bit quantization does, briefly."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=256)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

Transformers (AutoGPTQ backend)

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "gitarist/Qwen3-1.7B-GPTQ-Int4-g256-sym"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, device_map="auto")

vLLM

vllm serve gitarist/Qwen3-1.7B-GPTQ-Int4-g256-sym --quantization gptq

Evaluation

Perplexity on WikiText‑2 (raw, test split; 2048-token windows) and mean per‑token KL divergence of the quantized model's next‑token distribution against the original FP16 model, measured on the same text.

model WikiText‑2 PPL ΔPPL vs FP16 mean KL vs FP16
Qwen3‑1.7B (FP16, reference) 16.670 — 0.000
this model (Int4, g256, sym) 17.110 +0.440 0.273

Lower is better for all columns. The 4-bit model stays within ~0.44 perplexity of the FP16 baseline.

License

Inherits the license of the base model, Qwen/Qwen3-1.7B (Apache-2.0). This is a derivative quantized artifact.

Downloads last month
13
Safetensors
Model size
2B params
Tensor type
I32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gitarist/Qwen3-1.7B-GPTQ-Int4-g256-sym

Finetuned
Qwen/Qwen3-1.7B
Quantized
(410)
this model