Artemis 31B v1.2 — NVFP4A16 (weight-only)
Weight-only NVFP4 quant of TheDrummer/Artemis-31B-v1.2, a fine-tune of Gemma 4 31B, for vLLM on NVIDIA Blackwell (including DGX Spark / GB10).
- Scheme: LLM Compressor
NVFP4A16. Weights are FP4 with group size 16; activations stay BF16, so no calibration data was used. - Quantized: the 410
Linearlayers of the language model. - Left in BF16:
lm_head, embeddings, and the vision and audio towers. - Size: about 20 GB across two safetensors shards.
- Made with: llmcompressor 0.14.0, transformers 5.17.0, torch 2.13.0 (cu130). The recipe is in
recipe.yaml.
The chat template, tokenizer and processor config are copied unchanged from the source.
Why weight-only
The only other NVFP4 build of v1.2 quantizes activations to FP4 as well (W4A4). In vLLM 0.30.0 it inserted stray tokens mid-word in long roleplay replies ("de deciding", "deambles"), with or without speculative decoding. Replaying the same conversation against this build in vLLM 0.30.0 with MTP on, at temperature 0.94 and no top-k/top-p, produced none across four replies of about 450 tokens each.
Serving
vllm serve dgibbons/Artemis-31B-v1.2-NVFP4A16 --max-model-len 131072
Google's MTP drafter for stock Gemma 4 works as a speculator:
--speculative-config '{"method":"mtp","model":"google/gemma-4-31B-it-assistant","num_speculative_tokens":4}'
vLLM rejects min_p while speculative decoding is on; use top_k/top_p instead.
Attribution
The fine-tune is by TheDrummer, based on Google's Gemma 4 31B. Follow the terms of the source model and of Gemma 4. This repository contains only a precision conversion.
- Downloads last month
- 52