EXL3 quantization of Gemma-4-26B-A4B-StyleTune-V2, 4 bits per weight.

Use CPU offloading by setting EXL3_MOE_CPU_OFFLOAD.

E.g. with TabbyAPI:

EXL3_MOE_CPU_OFFLOAD=10 python main.py ...
Downloads last month
20
Safetensors
Model size
8B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for isogen/Gemma-4-26B-A4B-StyleTune-V2-exl3-4bpw

Quantized
(13)
this model