Instructions to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: llama cli -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: llama cli -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: ./llama-cli -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: ./build/bin/llama-cli -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Use Docker
docker model run hf.co/FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
- LM Studio
- Jan
- Ollama
How to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with Ollama:
ollama run hf.co/FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
- Unsloth Desktop
- Pi
How to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with Docker Model Runner:
docker model run hf.co/FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
- Lemonade
How to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Run and chat with the model
lemonade run user.Artemis-31B-v1.2-NVFP4-GGUF-NVFP4
List all available models
lemonade list
- Hermes Agent
How to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FaustianDeal/Artemis-31B-v1.2-NVFP4-GGUF:NVFP4" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Artemis 31B v1.2 — NVFP4 GGUF
This is a text-model GGUF conversion of TheDrummer/Artemis-31B-v1.2, a fine-tune of google/gemma-4-31B. The language model's eligible linear weights were calibrated and quantized to NVIDIA NVFP4, then repacked into GGUF. The original fine-tune is credited to TheDrummer; this repository contains a format and precision conversion.
| File | Size | SHA-256 |
|---|---|---|
Artemis-31B-v1.2-NVFP4.gguf |
19,313,589,984 bytes (18.0 GiB) | 0892d39cfb7295b07a8890da516cade061c6d4a9ca829bab846586fd831a7917 |
Use
Load Artemis-31B-v1.2-NVFP4.gguf as the model in a Gemma 4 and NVFP4-capable GGUF runtime. KoboldCpp 1.121 loaded this file and generated text successfully in a local test. The source model's Gemma 4 chat template is embedded in the GGUF.
Use the Gemma4 31B vision tower if you want image recognition.
Conversion details
- Source: TheDrummer/Artemis-31B-v1.2, revision
05d84790fceecefac4ee2adfb7cf33fdce2029f1. - Quantization: LLM Compressor's NVFP4 scheme, with 32 calibration samples of 2,048 tokens from
mit-han-lab/pile-val-backup. - Targets: eligible
Linearlayers. Vision and audio layers, embeddings, andlm_headwere excluded from NVFP4 quantization. - GGUF conversion:
llama.cppcommitb9ae43a5d4c27564963717281070991fa9b8c1bf, repacking the calibrated NVFP4 checkpoint without a second weight quantization pass. - File inspection: 1,653 tensors, including 410 NVFP4 tensors, 1,242 F32 tensors, and one BF16 tensor. The tokenizer and chat template match the source checkpoint by SHA-256.
Validation and limitations
On the 299-question ARC-Challenge validation split, using zero-shot, letter-only answers in KoboldCpp 1.121, NVFP4 scored 291/299 (97.3%) and a BF16 GGUF from the same v1.2 source scored 293/299 (98.0%). Both returned valid answers for every question; their predictions differed on two questions. This is one direct-answer benchmark run, so the small gap should not be treated as a general quality rating. The compressed-tensors checkpoint has not been benchmarked separately.
Attribution and terms
The original fine-tune is by TheDrummer, based on Gemma 4 31B. Follow the terms that apply to the source fine-tune and base model. The source repository did not declare a license in its model metadata when this card was prepared, so no license is asserted here.
- Downloads last month
- 747
4-bit