--- base_model: TheDrummer/Artemis-31B-v1.2 base_model_relation: quantized tags: - nvfp4 - compressed-tensors - gemma4 - vllm --- # Artemis 31B v1.2 — NVIDIA NVFP4 checkpoint This is a calibrated NVIDIA NVFP4 conversion of [TheDrummer/Artemis-31B-v1.2](https://huggingface.co/TheDrummer/Artemis-31B-v1.2), a fine-tune of [google/gemma-4-31B](https://huggingface.co/google/gemma-4-31B). It is provided as a **compressed-tensors safetensors checkpoint** for runtimes that support Gemma 4 and `nvfp4-pack-quantized`. The original fine-tune is credited to TheDrummer; this repository contains a precision conversion. The checkpoint has two safetensors shards totaling **20,446,828,448 bytes** (19.04 GiB), plus its configuration, tokenizer, chat template, processor configuration, recipe, and conversion manifest. Download the **whole repository** to use this format. The two weight shards alone are insufficient. ## Conversion details - Source: [TheDrummer/Artemis-31B-v1.2](https://huggingface.co/TheDrummer/Artemis-31B-v1.2), revision `05d84790fceecefac4ee2adfb7cf33fdce2029f1`. - Quantization: LLM Compressor's `NVFP4` scheme, with 32 calibration samples of 2,048 tokens from `mit-han-lab/pile-val-backup`. - Targets: eligible `Linear` layers. Vision and audio layers, embeddings, and `lm_head` were excluded from NVFP4 quantization. - Storage format: `compressed-tensors` with `nvfp4-pack-quantized` weights. The saved checkpoint contains 410 NVFP4 weight tensors. - The tokenizer and chat template match the source checkpoint by SHA-256. ## Validation and compatibility Both safetensors shards and the tensor index were inspected successfully; all 2,418 indexed tensors were found. **A generated response from this v1.2 checkpoint in vLLM or Transformers has not been verified.** Runtime compatibility, multimodal operation, and output quality should be tested in the target environment. The separately converted NVFP4 GGUF scored **291/299 (97.3%)** on a zero-shot direct-answer run of the [ARC-Challenge validation split](https://huggingface.co/datasets/allenai/ai2_arc), versus **293/299 (98.0%)** for a BF16 GGUF from the same v1.2 source. Only the GGUFs were run in KoboldCpp 1.121 for this comparison. Their results do not establish that this compressed-tensors checkpoint loads or generates correctly in another runtime. ## Attribution and terms The original fine-tune is by [TheDrummer](https://huggingface.co/TheDrummer), based on [Gemma 4 31B](https://huggingface.co/google/gemma-4-31B). Follow the terms that apply to the source fine-tune and base model. The source repository did not declare a license in its model metadata when this card was prepared, so no license is asserted here.