File size: 7,444 Bytes
9e3a41c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38bb75c
 
 
9e3a41c
 
 
 
38bb75c
 
 
 
 
 
9e3a41c
 
 
 
 
38bb75c
 
 
 
 
 
 
 
 
9e3a41c
 
 
 
38bb75c
9e3a41c
 
 
38bb75c
 
 
 
9e3a41c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
# Windows quick start: Ornith MTP OpenAI endpoint

This recipe reproduces the fast Windows laptop profile measured for
`ornith-1.0-35b-MTP-graft-down-Q4_0.gguf`:

- Windows 11 with Docker Desktop using WSL2;
- NVIDIA GPU visible to Docker (`docker run --rm --gpus all ...`);
- 8 GiB VRAM and approximately 44 GiB peak whole-system RAM in the measured
  RTX 4060 Laptop configuration;
- llama.cpp b10066 (`86a9c79f8`);
- dense and attention tensors on the GPU, MoE experts in host RAM;
- integrated MTP plus n-gram speculative decoding;
- OpenAI-compatible API on localhost only.

## 1. Prerequisites

Install or update:

1. NVIDIA Windows driver;
2. WSL2;
3. Docker Desktop with the WSL2 backend;
4. Git LFS or the Hugging Face CLI for the model download.

Confirm that Docker can see the GPU:

```powershell
docker run --rm --gpus all nvidia/cuda:12.8.1-base-ubuntu24.04 nvidia-smi
```

The validated 16K profile needs roughly 43–45 GiB of total RAM in use on the
test laptop. Close memory-heavy applications before starting on a 64 GiB
machine. An 8 GiB GPU is sufficient because `--cpu-moe` keeps the large expert
weights in host RAM.

## 2. Download the full MTP artifact

Create a model directory and an isolated virtual environment for the Hugging
Face CLI. Do not upgrade `huggingface_hub` in a shared Python installation:
new CLI releases can require a newer `click` than packages such as `gTTS`.

```powershell
New-Item -ItemType Directory -Force C:\Models\Ornith-MTP | Out-Null

$HfVenv = 'C:\Models\Ornith-MTP\.hf-cli'
py -m venv $HfVenv
& "$HfVenv\Scripts\python.exe" -m pip install --upgrade pip
& "$HfVenv\Scripts\python.exe" -m pip install "huggingface_hub[hf_xet]==1.24.0"

& "$HfVenv\Scripts\hf.exe" download `
  petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF `
  ornith-1.0-35b-MTP-graft-down-Q4_0.gguf `
  --local-dir C:\Models\Ornith-MTP
```

The download is approximately 19.4 GiB. Xet may spend time scanning chunks
without continuously printing progress; leave the command running until the
PowerShell prompt returns. A partially downloaded file is resumed on the next
identical command.

This creates a real standalone copy downloaded from Hugging Face. The recipe
does not use a hard link, symbolic link, or an already installed LM Studio
model.

Expected file:

```text
C:\Models\Ornith-MTP\ornith-1.0-35b-MTP-graft-down-Q4_0.gguf
Size: 20,329,342,112 bytes (18.933 GiB)
SHA-256: 365a7c02dfd320b9696f189d6dc12bd2b0eabb9f8e58ba9fc8cab3af93c0234b
```

This clean download and endpoint flow was re-tested on Windows on 19 July
2026, including checksum verification, Docker startup, `/health`, `/v1/models`,
and `/v1/chat/completions`.

Verify it:

```powershell
(Get-FileHash `
  C:\Models\Ornith-MTP\ornith-1.0-35b-MTP-graft-down-Q4_0.gguf `
  -Algorithm SHA256).Hash.ToLowerInvariant()
```

Do not use the separate `LM Studio compatible, no integrated MTP` file for
this endpoint. It deliberately omits the speculative prediction layer.

## 3. Start the endpoint

The repository includes `run-ornith-mtp-windows.ps1`. Download it beside the
model card or copy it from this repository, then run:

```powershell
Set-ExecutionPolicy -Scope Process Bypass
.\run-ornith-mtp-windows.ps1 `
  -ModelDir C:\Models\Ornith-MTP `
  -Port 18081
```

The script pins the exact tested image:

```text
ghcr.io/ggml-org/llama.cpp@sha256:1b3d1458ccda7287feab41b8001311acc03e24cde99ec0a2908fe83830562f38
llama.cpp b10066 (86a9c79f8)
```

Equivalent direct command:

```powershell
$Image = 'ghcr.io/ggml-org/llama.cpp@sha256:1b3d1458ccda7287feab41b8001311acc03e24cde99ec0a2908fe83830562f38'
$ModelDir = 'C:\Models\Ornith-MTP'

docker run --detach --rm `
  --name ornith-mtp-api `
  --gpus all `
  --publish 127.0.0.1:18081:8080 `
  --mount "type=bind,source=$ModelDir,target=/models,readonly" `
  $Image `
  --model /models/ornith-1.0-35b-MTP-graft-down-Q4_0.gguf `
  --alias ornith-1.0-35b-mtp `
  --host 0.0.0.0 --port 8080 `
  --ctx-size 16384 --parallel 1 `
  --n-gpu-layers all --cpu-moe --no-mmap `
  --batch-size 2048 --ubatch-size 512 `
  --flash-attn on `
  --cache-type-k q8_0 --cache-type-v q8_0 `
  --jinja --metrics `
  --spec-type ngram-mod,draft-mtp `
  --spec-draft-n-max 2 `
  --spec-draft-n-min 1 `
  --spec-draft-p-min 0.20 `
  --spec-ngram-mod-n-min 48 `
  --spec-ngram-mod-n-max 64 `
  --spec-ngram-mod-n-match 24
```

The host publish address is deliberately `127.0.0.1`, not `0.0.0.0`. This
prevents unauthenticated access from the LAN. Put nginx/Caddy with TLS and an
API key in front of llama-server before exposing it to other machines.

## 4. Check readiness and API

Model loading can take tens of seconds:

```powershell
docker logs -f ornith-mtp-api
```

In another terminal:

```powershell
Invoke-RestMethod http://127.0.0.1:18081/health
Invoke-RestMethod http://127.0.0.1:18081/v1/models
```

OpenAI-compatible smoke request:

```powershell
$Body = @{
  model = 'ornith-1.0-35b-mtp'
  messages = @(
    @{ role = 'user'; content = 'Reply with exactly: READY' }
  )
  temperature = 0
  max_tokens = 64
} | ConvertTo-Json -Depth 8

$Response = Invoke-RestMethod `
  -Method Post `
  -Uri http://127.0.0.1:18081/v1/chat/completions `
  -ContentType 'application/json' `
  -Body $Body

$Response.choices[0].message
```

Use these provider values in Orion or another OpenAI client:

```text
Base URL: http://127.0.0.1:18081/v1
API key:  local-only-placeholder
Model:    ornith-1.0-35b-mtp
```

llama-server does not require the placeholder key in this localhost profile,
but many OpenAI clients require a non-empty string in their configuration.

## 5. Observe and stop

```powershell
# llama.cpp request and MTP statistics
Invoke-WebRequest http://127.0.0.1:18081/metrics -UseBasicParsing |
  Select-Object -ExpandProperty Content

# GPU telemetry
nvidia-smi `
  --query-gpu=temperature.gpu,utilization.gpu,memory.used,power.draw `
  --format=csv,noheader,nounits

# logs
docker logs --tail 100 ornith-mtp-api

# stop and release RAM/VRAM
docker stop ornith-mtp-api
```

## Expected performance on the measured laptop

RTX 4060 Laptop 8 GiB, i7-13650HX and 64 GiB RAM:

| Workload | Measured decode |
|---|---:|
| 1K prompt, ordinary output | 28.74 tok/s |
| 8K prompt, high MTP acceptance | 66.88 tok/s |
| repeated 8K code | 67.23 tok/s |

The 60–67 tok/s range requires high speculative acceptance. Ordinary agent
traffic should not be advertised as a fixed 67 tok/s minimum. The same hybrid
without speculation measured about 27–30 tok/s.

## Common problems

### `could not select device driver ... gpu`

Update Docker Desktop and the NVIDIA driver, ensure Docker uses the WSL2
backend, restart Docker Desktop, then repeat the CUDA `nvidia-smi` check.

### Container exits during model loading

Check `docker logs ornith-mtp-api`, close memory-heavy programs, and confirm at
least about 45 GiB of RAM is available for the measured profile.

### Only 1–2 tok/s

Do not replace `--cpu-moe --n-gpu-layers all` with an arbitrary partial layer
split. The rejected 15-layer split forced expert traffic across CPU/GPU memory
and reduced realistic prompt decode to about 1.2 tok/s.

### MTP flags are unknown

The image is too old or a different runtime is being used. Run:

```powershell
docker run --rm `
  ghcr.io/ggml-org/llama.cpp@sha256:1b3d1458ccda7287feab41b8001311acc03e24cde99ec0a2908fe83830562f38 `
  --version
```

It must report llama.cpp version 10066 (`86a9c79f8`) for this exact recipe.