Artemis 31B v1.2 — NVFP4 GGUF

This is a text-model GGUF conversion of TheDrummer/Artemis-31B-v1.2, a fine-tune of google/gemma-4-31B. The language model's eligible linear weights were calibrated and quantized to NVIDIA NVFP4, then repacked into GGUF. The original fine-tune is credited to TheDrummer; this repository contains a format and precision conversion.

File Size SHA-256
Artemis-31B-v1.2-NVFP4.gguf 19,313,589,984 bytes (18.0 GiB) 0892d39cfb7295b07a8890da516cade061c6d4a9ca829bab846586fd831a7917

Use

Load Artemis-31B-v1.2-NVFP4.gguf as the model in a Gemma 4 and NVFP4-capable GGUF runtime. KoboldCpp 1.121 loaded this file and generated text successfully in a local test. The source model's Gemma 4 chat template is embedded in the GGUF.

Use the Gemma4 31B vision tower if you want image recognition.

Conversion details

  • Source: TheDrummer/Artemis-31B-v1.2, revision 05d84790fceecefac4ee2adfb7cf33fdce2029f1.
  • Quantization: LLM Compressor's NVFP4 scheme, with 32 calibration samples of 2,048 tokens from mit-han-lab/pile-val-backup.
  • Targets: eligible Linear layers. Vision and audio layers, embeddings, and lm_head were excluded from NVFP4 quantization.
  • GGUF conversion: llama.cpp commit b9ae43a5d4c27564963717281070991fa9b8c1bf, repacking the calibrated NVFP4 checkpoint without a second weight quantization pass.
  • File inspection: 1,653 tensors, including 410 NVFP4 tensors, 1,242 F32 tensors, and one BF16 tensor. The tokenizer and chat template match the source checkpoint by SHA-256.

Validation and limitations

On the 299-question ARC-Challenge validation split, using zero-shot, letter-only answers in KoboldCpp 1.121, NVFP4 scored 291/299 (97.3%) and a BF16 GGUF from the same v1.2 source scored 293/299 (98.0%). Both returned valid answers for every question; their predictions differed on two questions. This is one direct-answer benchmark run, so the small gap should not be treated as a general quality rating. The compressed-tensors checkpoint has not been benchmarked separately.

Attribution and terms

The original fine-tune is by TheDrummer, based on Gemma 4 31B. Follow the terms that apply to the source fine-tune and base model. The source repository did not declare a license in its model metadata when this card was prepared, so no license is asserted here.

Downloads last month
747
GGUF
Model size
31B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF

Quantized
(20)
this model