Artemis 31B v1.2 — NVFP4A16 (weight-only)

Weight-only NVFP4 quant of TheDrummer/Artemis-31B-v1.2, a fine-tune of Gemma 4 31B, for vLLM on NVIDIA Blackwell (including DGX Spark / GB10).

  • Scheme: LLM Compressor NVFP4A16. Weights are FP4 with group size 16; activations stay BF16, so no calibration data was used.
  • Quantized: the 410 Linear layers of the language model.
  • Left in BF16: lm_head, embeddings, and the vision and audio towers.
  • Size: about 20 GB across two safetensors shards.
  • Made with: llmcompressor 0.14.0, transformers 5.17.0, torch 2.13.0 (cu130). The recipe is in recipe.yaml.

The chat template, tokenizer and processor config are copied unchanged from the source.

Why weight-only

The only other NVFP4 build of v1.2 quantizes activations to FP4 as well (W4A4). In vLLM 0.30.0 it inserted stray tokens mid-word in long roleplay replies ("de deciding", "deambles"), with or without speculative decoding. Replaying the same conversation against this build in vLLM 0.30.0 with MTP on, at temperature 0.94 and no top-k/top-p, produced none across four replies of about 450 tokens each.

Serving

vllm serve dgibbons/Artemis-31B-v1.2-NVFP4A16 --max-model-len 131072

Google's MTP drafter for stock Gemma 4 works as a speculator:

--speculative-config '{"method":"mtp","model":"google/gemma-4-31B-it-assistant","num_speculative_tokens":4}'

vLLM rejects min_p while speculative decoding is on; use top_k/top_p instead.

Attribution

The fine-tune is by TheDrummer, based on Google's Gemma 4 31B. Follow the terms of the source model and of Gemma 4. This repository contains only a precision conversion.

Downloads last month
52
Safetensors
Model size
18B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dgibbons/Artemis-31B-v1.2-NVFP4A16

Quantized
(20)
this model