Gemma-3-1B-it GPTQ INT4/G128 Safetensors

Pre-LLiMa Hugging Face checkpoint for Sima.ai compilation.

Quantization

  • All language Linear layers and lm_head: GPTQ, symmetric INT4, group size 128, static actorder.
  • Group size 128 is required because 1152-wide layers do not support G256.
  • Calibration: HuggingFaceH4/ultrachat_200k, 512 samples, 1024 tokens, batch size 1.

Full WikiText evaluation

EleutherAI/wikitext_document_level, wikitext-2-raw-v1, full matched run (2026-07-16).

Checkpoint Word perplexity
google/gemma-3-1b-it 35.970619
This GPTQ checkpoint 42.386377

Reproduce

python quantize.py --model-path /project/mlasw/share/huggingface/models--google--gemma-3-1b-it --output-dir /path/to/output

Validation and limitations

Saved scales were checked for finite values. Transformers load/generation smoke test is pending. Do not publish this artifact without resolving its quality regression and recording the smoke test.

Downloads last month
79
Safetensors
Model size
1B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for simaai/Gemma-3-1B-it-GPTQ-Safetensors

Quantized
(457)
this model

Collection including simaai/Gemma-3-1B-it-GPTQ-Safetensors