Artemis 31B v1.2 | TheDrummer | AWQ INT4 | OpenVINO
Original fine-tune by TheDrummer. OpenVINO conversion by Wondernutts.
This is the OpenVINO INT4 conversion of TheDrummer/Artemis-31B-v1.2, built on Google's dense Gemma 4 31B architecture. This is a compressed deployment derivative, not a new fine-tune.
Why bring Artemis to OpenVINO?
Big credit to TheDrummer for the creative-writing and roleplay work behind these releases. His models bring welcome variety to the RP community, and making that work accessible to Intel Arc users is exactly why this conversion exists. If you enjoy Artemis, please support and credit the original creator and explore TheDrummer's other models.
The expressive, dynamic feel of TheDrummer's Orion was what drew me toward Artemis next. That is a reason to explore this release, not a claim that this Artemis conversion has already won a quality comparison.
At a glance
| Item | This release |
|---|---|
| Model | Artemis 31B v1.2, dense Gemma 4; not the 26B-A4B MoE |
| Language weights | Asymmetric INT4 AWQ, group size 128, ratio 1.0 |
| Calibration | Data-free AWQ; no supplied conversation or benchmark dataset |
| Repository payload | Approximately 18.15 GiB / 19.49 GB at release |
| RoPE | Four F32 lookup tables, positions 0 through 131071 |
| Intended deployment | Intel Arc on Linux with a compatible OpenVINO GenAI runtime |
| Reference development hardware | Intel Arc Pro B70, 32 GB VRAM, Linux |
| Release status | Conversion and structural checks passed; Artemis-specific inference qualification pending |
Repository size is not total VRAM usage. KV cache, graph compilation, temporary buffers, and vision add memory requirements. No 16 GB GPU fit or offload capability is promised. Even a 32 GB GPU needs an appropriate context and memory budget. The 131K LUT capacity is a positional bound, not a tested usable context window.
TheDrummer's sampling defaults
The shipped generation_config.json contains:
| Setting | Value |
|---|---|
do_sample |
true |
temperature |
1.0 |
top_k |
64 |
top_p |
0.95 |
Repetition, frequency, presence, and Min-P values are not specified in that file. Do not attribute a frontend's additional defaults to TheDrummer.
The upstream card supports both reasoning-enabled and direct-answer use with the Gemma 4 template, and recommends experimenting with sampling conservatively. It also links community sampler recommendations. Keep the provided tokenizer and chat template; do not substitute a generic ChatML template.
Running with OpenVINO GenAI
This is OpenVINO IR, not GGUF and not a Transformers weight checkpoint. Transformers is used below only to format the prompt with the supplied tokenizer. Inference runs through OpenVINO GenAI.
Use a separate environment and a compatible Intel GPU driver. The following is a starting example, not an Artemis-tested runtime certification:
python -m pip install "openvino-genai==2026.3.0.0" "transformers==5.5.0" huggingface_hub
import openvino_genai as ov_genai
from huggingface_hub import snapshot_download
from transformers import AutoTokenizer
model_dir = snapshot_download("Wondernutts/Artemis-31B-v1.2-int4-ov")
tokenizer = AutoTokenizer.from_pretrained(model_dir)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "You are a Whiterun innkeeper. Greet a traveler in four sentences."}],
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
pipe = ov_genai.VLMPipeline(model_dir, "GPU")
config = ov_genai.GenerationConfig()
config.max_new_tokens = 256
config.do_sample = True
config.temperature = 1.0
config.top_k = 64
config.top_p = 0.95
config.repetition_penalty = 1.0 # Neutral example setting, not an upstream recommendation.
config.apply_chat_template = False # Already formatted above; do not apply twice.
print(pipe.generate(prompt, generation_config=config))
Start with a short prompt before increasing context. Changing enable_thinking to True requests reasoning and requires more output budget; applications should separate reasoning from the displayed answer. The example does not configure a production scheduler, prefix caching, quantized KV cache, or concurrency limits.
Vision embedding artifacts are included, but image/video behavior has not been qualified for this conversion. No audio inference is claimed. This text example does not validate multimodal support.
Wondernutts' custom OpenVINO work
The custom runtime kernels are separate software, not baked into the weights in this repository. Installing the pip packages above does not install that fork. Follow its documented build and model-compatibility instructions for custom serving; a 12B or 26B example is not automatically a validated dense-31B profile.
No Artemis-specific prompt-processing or decode benchmark is published yet. Numbers from Orion, Chimera-X, or another 31B checkpoint should not be presented as measurements of this model. The B70/Linux development setup is not a claim of equivalent performance on B50, A770, Windows, or CPU.
Conversion and checks
- Source revision:
05d84790fceecefac4ee2adfb7cf33fdce2029f1. - Data-free AWQ INT4 asymmetric g128; no scale estimation, GPTQ, or RP-log calibration.
- No MoE router exclusion: this checkpoint is dense.
- Source tokenizer/chat template retained; OpenVINO tokenizer and detokenizer included.
- Packed INT4/INT8 constant signatures, including their bytes, checked unchanged across the LUT patch.
- Four F32 LUT tables verified, with a position clamp of 131071 and no remaining runtime Sin/Cos in the language graph.
- LUT toolkit revision:
546089b1628c525357245e82079159a3a75973b7.
Conversion packages: OpenVINO 2026.3.0, OpenVINO Tokenizers 2026.3.0.0, NNCF 3.3.0, Transformers 5.5.0, Optimum 2.3.0, Optimum Intel 2.2.0.dev0+8491053, CPU Torch 2.11.0+cpu.
See conversion metadata and structural verification record. These checks establish the conversion structure, not parity with full precision, reliable long-context coherence, or an OOM-free serving configuration. B70 inference/coherence and multimodal validation remain pending.
Credits and terms
- TheDrummer: Artemis fine-tune and original release.
- Google: Gemma 4 base model architecture and weights.
- Wondernutts: this OpenVINO conversion, LUT integration, and custom runtime work linked above.
- OpenVINO, NNCF, Optimum Intel and Hugging Face contributors: conversion and inference tooling.
The upstream Artemis card does not declare a separate license identifier. Google's base-model metadata identifies Apache-2.0 and links its Gemma 4 license information. Review the original Artemis repository, the preserved upstream model card, and applicable base-model terms before redistribution or commercial use. This conversion does not replace or waive upstream terms.
- Downloads last month
- 32