Instructions to use stuchapin/Carnice-V2-27B-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Use Docker
docker model run hf.co/stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "stuchapin/Carnice-V2-27B-MTP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "stuchapin/Carnice-V2-27B-MTP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
- Ollama
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with Ollama:
ollama run hf.co/stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with Docker Model Runner:
docker model run hf.co/stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
- Lemonade
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Carnice-V2-27B-MTP-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use stuchapin/Carnice-V2-27B-MTP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "stuchapin/Carnice-V2-27B-MTP-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Carnice-V2-27B + MTP — GGUF for Ampere/older NVIDIA
GGUF builds of kai-os/Carnice-V2-27b with the MTP (Multi-Token Prediction) speculative-decoding head grafted in, packaged for llama.cpp. The MTP tensors come from sakamakismile/Carnice-V2-27b-NVFP4-TEXT-MTP (whose NVFP4 build is Blackwell-only). This repo makes the same MTP-accelerated decoding available on Ampere / Ada / pre-Blackwell GPUs via GGUF.
⚠️ Requires unmerged llama.cpp PR #22673
These GGUFs only work with llama.cpp PR #22673, which adds MTP support and is still in draft as of 2026-05-11. Mainline llama.cpp will fail to load them.
git clone https://github.com/ggml-org/llama.cpp.git && cd llama.cpp
git fetch origin pull/22673/head:mtp-pr && git checkout mtp-pr
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release -DCMAKE_CUDA_ARCHITECTURES=86
cmake --build build --target llama-server llama-quantize -j8
(CUDA_ARCHITECTURES=86 for RTX 3090; 89 for 4090/Ada, 90 for Hopper.)
Variants
| File | Base quant | Size | Quality | Best for |
|---|---|---|---|---|
Carnice-V2-27B-Q8_0-mtp.gguf |
Q8_0 | 28 GB | A+ (near-lossless) | Max quality, 32GB+ VRAM or multi-GPU |
Carnice-V2-27B-Q6_K-mtp.gguf |
Q6_K | 22 GB | A+ | Best mainstream quality |
Carnice-V2-27B-Q5_K_M-mtp.gguf |
Q5_K_M | 19 GB | A | Single 24GB card, A-tier quality |
Carnice-V2-27B-Q4_K_M-mtp.gguf |
Q4_K_M | 17 GB | B+ | Speed/quality balance |
Carnice-V2-27B-IQ4_XS-mtp.gguf |
IQ4_XS | 15 GB | B+ | Speed leader on Ampere |
All variants keep the MTP block (blk.64.*) at BF16. Quantizing the MTP head didn't help speed and risked acceptance-rate degradation in testing.
Benchmarks — real agent workload
Measured on 2× RTX 3090 (Q8_0 and Q6_K via tensor-split, others single-GPU), --spec-draft-n-max 1, -c 32768, KV q4_0, flash attention on, thinking mode enabled. Tested with 7 representative prompts from a Hermes-style agent workload:
- tool-call-meeting — emit JSON tool calls for a CRM workflow
- morning-brief-synthesis — 5-bullet summary from agent context
- concierge-permit-extract — structured extraction from permit text
- draft-touch-card — short-form professional writing
- lease-clause-explain — domain knowledge / explanation
- long-context-summarize — 30-lead CRM dump → top 3
- reasoning-lease-vs-buy — multi-step numerical reasoning
Aggregate
| Variant | Mean tok/s | MTP acceptance | Total tokens / 7 prompts |
|---|---|---|---|
| Q8_0 | 34.0 | 82% | 12,144 |
| Q6_K | 38.4 | 84% | 12,680 |
| Q5_K_M | 39.9 | 83% | 12,660 |
| Q4_K_M | 44.0 | 81% | 12,084 |
| IQ4_XS | 47.7 | 82% | 13,111 |
Per-prompt tok/s
| Prompt | Q8_0 | Q6_K | Q5_K_M | Q4_K_M | IQ4_XS |
|---|---|---|---|---|---|
| tool-call-meeting | 28.7 | 26.9 | 32.5 | 38.3 | 35.0 |
| morning-brief-synthesis | 34.9 | 39.6 | 42.9 | 46.8 | 49.4 |
| concierge-permit-extract | 33.8 | 38.4 | 40.0 | 44.8 | 47.6 |
| draft-touch-card | 34.3 | 38.2 | 41.2 | 45.2 | 48.2 |
| lease-clause-explain | 33.7 | 38.9 | 39.9 | 44.6 | 48.7 |
| long-context-summarize | 34.7 | 38.9 | 40.0 | 44.6 | 48.0 |
| reasoning-lease-vs-buy | 33.2 | 37.6 | 38.7 | 41.9 | 46.6 |
MTP acceptance is steady 81–84% across all variants — quant choice didn't degrade the draft head's accuracy. Speed differences come almost entirely from the base model's memory bandwidth per token.
Reference: non-MTP vanilla decode on the same hardware sits at ~35 tok/s, so even the slowest MTP variant (Q8_0) is roughly even with vanilla; IQ4_XS gives ~1.4× over the vanilla baseline on this same workload mix.
Method
The whole thing is reproducible end-to-end. Headline steps:
1. Source materials
- Base weights:
kai-os/Carnice-V2-27b(BF16 safetensors, 55 GB) — the Hermes-style SFT of Qwen3.6-27B. - MTP head: 15
mtp.*BF16 tensors (~850 MB) extracted fromsakamakismile/Carnice-V2-27b-NVFP4-TEXT-MTP. sakamakismile kept these tensors unquantized inside their NVFP4 file (the surrounding model is FP8/U8), so no dequantization required — just lift them out withsafetensors.safe_open.
2. Merge
A small Python script symlinks the base BF16 shards into a working directory, writes the 15 MTP tensors as a new model-mtp.safetensors, regenerates model.safetensors.index.json, and rewrites config.json to set language_model_only: true (Carnice's base config carries a vision tower we don't ship).
3. Convert to GGUF
convert_hf_to_gguf.py from llama.cpp PR #22673 has first-class support for Qwen3_5ForConditionalGeneration and automatically remaps the mtp.* tensors to llama.cpp's blk.64.nextn.* naming. Output is a single BF16 GGUF (~55 GB).
4. Quantize
llama-quantize from the same PR build, with --tensor-type "blk\.64\..*=bf16" to keep the MTP block at full precision. Five targets: Q8_0, Q6_K, Q5_K_M, Q4_K_M, IQ4_XS.
5. The critical tuning find — --spec-draft-n-max 1
Default community guides recommend --spec-draft-n-max 3. For Carnice this is wrong by a wide margin. The MTP head has mtp_num_hidden_layers: 1 — it was trained to predict exactly one token ahead. Drafting more tokens recurses the same head autoregressively; errors cascade rapidly.
We swept --spec-draft-n-max values 1, 2, 3, 5, 7 on the same workload:
--spec-draft-n-max |
Mean tok/s | Notes |
|---|---|---|
| 1 | 52.6 | Optimal — acceptance stays 80%+ |
| 2 | (crash on our hardware) | |
| 3 | 46.6 | Acceptance drops to 26% on creative outputs |
| 5 | 38.5 | Marginal acceptance, draft cost dominates |
| 7 | (crashed) |
General rule: set --spec-draft-n-max = mtp_num_hidden_layers from the model's config.
6. What didn't help
We also exhaustively swept (each at --spec-draft-n-max 1):
- MTP block quant (BF16 / Q8_0 / Q5_K) — all within ~2% noise. Default to BF16 for safety.
- KV cache type (q4_0 / q8_0) — within noise; f16 OOMs at 64K context.
- Batch / ubatch sizes — within noise.
- Thread counts (1, 4, 8, auto) — within noise (the workload is GPU-bound).
- CUDA graph disable — flag absent in PR build.
ik_llama.cppfork — ~40% slower for this architecture. Carnice has 48 linear-attention (Mamba2-style) layers; mainline+PR-22673's kernels for those are noticeably more optimized than ik_llama.cpp's. Stick with mainline+PR.
Running
Single-GPU (≤19 GB models)
./build/bin/llama-server \
-m Carnice-V2-27B-IQ4_XS-mtp.gguf \
--alias carnice-v2-mtp \
-c 65536 --parallel 1 -ngl 99 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--flash-attn on --no-context-shift \
--jinja --reasoning-format deepseek --reasoning-budget 4096 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0 \
--batch-size 2048 --ubatch-size 1024 --metrics \
--spec-type mtp --spec-draft-n-max 1 \
--fit off \
--host 127.0.0.1 --port 8001
Dual-GPU (Q6_K / Q8_0)
Add --tensor-split 14,14 and CUDA_VISIBLE_DEVICES=0,1. Expect ~15% PCIe overhead vs a single ≥32GB card.
Verify MTP is active
In the server startup log:
set_mtp: MTP draft head registered (ctx_mtp=..., n_embd=5120)
slot load_model: speculative decoding context initialized
In per-request response timings:
"timings": {
"predicted_per_second": 47.6,
"draft_n": 13,
"draft_n_accepted": 13
}
draft_n_accepted / draft_n should be ≥ 70% — we routinely see 100% on tool-call workloads.
Build recipe (reproduce from source)
# Merge MTP tensors into BF16 base
import json
from pathlib import Path
from safetensors import safe_open
from safetensors.torch import save_file
BASE = Path("./kai-os--Carnice-V2-27b")
MTP_SRC = Path("./sakamakismile--Carnice-V2-27b-NVFP4-TEXT-MTP")
OUT = Path("./merged")
OUT.mkdir(exist_ok=True)
for f in BASE.iterdir():
if f.name in ("model.safetensors.index.json", "config.json"): continue
(OUT / f.name).symlink_to(f.resolve())
mtp = {}
with safe_open(MTP_SRC / "model.safetensors", framework="pt") as fp:
for k in fp.keys():
if k.startswith("mtp."):
mtp[k] = fp.get_tensor(k)
save_file(mtp, str(OUT / "model-mtp.safetensors"))
with open(BASE / "model.safetensors.index.json") as f: idx = json.load(f)
for k in mtp: idx["weight_map"][k] = "model-mtp.safetensors"
with open(OUT / "model.safetensors.index.json", "w") as f: json.dump(idx, f)
with open(BASE / "config.json") as f: cfg = json.load(f)
cfg["language_model_only"] = True
for k in ("vision_config", "image_token_id", "video_token_id"):
cfg.pop(k, None)
with open(OUT / "config.json", "w") as f: json.dump(cfg, f, indent=2)
# Convert merged HF checkpoint to BF16 GGUF
python convert_hf_to_gguf.py ./merged \
--outfile Carnice-V2-27B-bf16-mtp.gguf --outtype bf16
# Quantize, preserving the MTP block at BF16
./build/bin/llama-quantize --tensor-type "blk\.64\..*=bf16" \
Carnice-V2-27B-bf16-mtp.gguf Carnice-V2-27B-IQ4_XS-mtp.gguf IQ4_XS
Known issues
- PR #22673 trips
free(): invalid pointerSIGABRT on graceful shutdown. Runtime is clean; the kernel reclaims the process. Cosmetic. Likely fixed before PR merges. - Vision input + MTP crashes in PR #22673. These GGUFs are text-only (vision config stripped).
- Prefill is ~half-speed with MTP enabled vs non-MTP. Matters most for very long prompts; decode is where MTP wins back.
License & attribution
This work is Apache 2.0 (inherits from kai-os/Carnice-V2-27b). The underlying Qwen3.6 base is subject to the Tongyi Qianwen License — commercial use exceeding 700K MAU requires a separate license from Alibaba.
Credit chain
- Qwen team (Alibaba) — Qwen3.6 base model
- kai-os — Carnice-V2-27B SFT (Hermes-style agent training)
- sakamakismile — MTP head graft + NVFP4 packaging; we extracted the MTP tensors from their build
- am17an and llama.cpp contributors — MTP support via PR #22673
- stuchapin (this repo) — GGUF + MTP fusion for Ampere,
--spec-draft-n-max 1tuning finding, end-to-end agent-workload benchmarks
Changelog
- 2026-05-11 — Initial release. 5 GGUF variants (Q8_0, Q6_K, Q5_K_M, Q4_K_M, IQ4_XS), all clean from BF16 source, MTP block preserved at BF16. llama.cpp PR #22673 commit
5d5f1b4. Real-workload benchmarks across 7 Hermes-style agent prompts.
- Downloads last month
- 108
4-bit
5-bit
6-bit
8-bit
docker model run hf.co/stuchapin/Carnice-V2-27B-MTP-GGUF: