Llamacpp imatrix special quantizations of Artemis-31B-v1.2

This repo contains special long-context quants of TheDrummer/Artemis-31B-v1.2 - see original model page for details on how to use this model.

Made with llama.cpp version b11211.

Quantization details

I used bartowski's imatrix file for this model.

The quants here are specialized for long context. The following layers are preserved with overridden, higher quantization level:

  • input/output layer (token_embed - only one layer here since gemma 4 uses tied embeddings)
  • attention layers - gemma 4 uses sliding-window and global attention in ratio 5:1, so global attention is the primary focus - global attention layers are always in bf16 and other attention layers have lower quantization level

Quants are focusing on 32GB VRAM setup, so the base quantization level is Q5_K_M + overrides.

Available quants:

  • Artemis-31B-v1.2-Q5_K_M_hb8-ga8-a6-fl.gguf - base level Q5_K_M, input/output Q8_0, global attention Q8_0, other attention Q6_K, ffn_down Q6_K, layers 0, 1 and 59 minimum at Q6_K - 6.31 BPW
Downloads last month
189
GGUF
Model size
31B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DeusImperator/Artemis-31B-v1.2-GGUF-long-ctx

Quantized
(20)
this model