Instructions to use simaai/Gemma-4-E2B-it-GPTQ-Safetensors with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use simaai/Gemma-4-E2B-it-GPTQ-Safetensors with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("simaai/Gemma-4-E2B-it-GPTQ-Safetensors") model = AutoModelForMultimodalLM.from_pretrained("simaai/Gemma-4-E2B-it-GPTQ-Safetensors", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Gemma-4-E2B-it GPTQ
This is a local post-training quantized checkpoint of google/gemma-4-E2B-it for later LLiMa compilation and deployment on Sima.ai hardware. It must not be uploaded as part of the current VLM batch.
Source revision: Not captured in the flattened local source directory. An immutable upstream revision must be recovered and recorded before any release or redistribution.
This checkpoint remains subject to the upstream Apache-2.0 license, intended-use guidance, and limitations.
Quantization
| Component | Method | Weight format | Effective targets |
|---|---|---|---|
Decoder Linear layers and lm_head |
GPTQ | symmetric INT4, group size 256, static act-order | 277 |
| Vision encoder Linear layers | GPTQ | symmetric INT8, per-channel, static act-order | 113 |
| Audio tower | none | source BF16 | 134 explicit exceptions |
| Vision/audio embedding projectors | none | source BF16 | 2 explicit exceptions |
The lm_head follows the default GPTQ INT4/G256 policy; no INT8 escalation was made. Calibration used lmms-lab/flickr30k test[:512], deterministic dataset order, 512 image-text samples, sequence length 2048, batch size 1, no concatenation, and no padding to maximum length. Exact targets and exceptions are saved in recipe.yaml.
Evaluation
Full MMStar used all 1,500 examples, VLMEvalKit commit 7055d3010c38ccb5dcae1bc9535ca19c7fe5d79f, deterministic generation, a matched letter-only system prompt, maximum 16 new tokens, SDPA, and exact_matching without an API judge.
| Checkpoint | Overall accuracy |
|---|---|
| Source | 35.6667% |
| Quantized | 29.7333% |
| Absolute change | -5.9333 percentage points |
| Relative change | -16.6355% |
Validation: MMStar inference passed (1,500/1,500). Raw results are under vlmeval_results/mmstar_full/.
Full matched WikiText-2 word perplexity (EleutherAI/wikitext_document_level, wikitext-2-raw-v1, no example limit; 2026-07-18): source 3362.9125; this checkpoint 618.3263; raw-text delta -2744.5862 (-81.61%). Independent native-logit scoring confirmed that the source's untemplated raw-text likelihood is intrinsically anomalous, so this lower PPL is not a quality improvement and must not be used as a release gate. It conflicts with the full MMStar regression. Investigation and raw evidence: releases/vlm-full-wikitext-perplexity/GEMMA4_PPL_INVESTIGATION.md and perplexity_results/vlm_full_wikitext/.
Reproduction
This directory contains the exact quantize.py, recipe.yaml, and versions.txt used for this artifact:
python quantize.py \
--model-path /path/to/source-model \
--output-dir /path/to/new-quantized-model
Environment
Python: 3.13.2
torch: 2.11.0+cu128
CUDA: 12.8
transformers: 5.10.1
llm-compressor: 0.12.0
auto-round: 0.13.0
compressed-tensors: 0.17.1
Deployment
This is the pre-LLiMa quantized Hugging Face artifact. LLiMa compiler output must remain separate unless its format and redistribution are approved. No Hugging Face upload is authorized for this batch.
Limitations
Quantization quality can vary by language, visual domain, prompt format, context length, and downstream runtime. MMStar does not replace application-specific validation. Audio weights and modality projectors remain in source precision.
- Downloads last month
- 33