petr567 commited on
Commit
38bb75c
·
verified ·
1 Parent(s): 9e3a41c

docs: verify clean Windows download and endpoint recipe

Browse files
Files changed (2) hide show
  1. README.md +5 -0
  2. WINDOWS_ENDPOINT_QUICKSTART.md +23 -5
README.md CHANGED
@@ -47,6 +47,11 @@ the measured CPU-MoE/GPU-dense MTP profile.
47
  Endpoint: `http://127.0.0.1:18081/v1`. Do not substitute the no-MTP LM Studio
48
  file if speculative acceleration is required.
49
 
 
 
 
 
 
50
  This repository contains a single, deployment-ready GGUF optimized and tested
51
  for batch-1 agent workloads on AMD Ryzen AI MAX+ 395 / Radeon 8060S (`gfx1151`,
52
  128 GiB UMA) with the Vulkan backend of llama.cpp b9994. The same artifact was
 
47
  Endpoint: `http://127.0.0.1:18081/v1`. Do not substitute the no-MTP LM Studio
48
  file if speculative acceleration is required.
49
 
50
+ The documented flow was re-tested from a fresh `C:\Models\Ornith-MTP`
51
+ directory using a real Hugging Face download (no hard links), SHA-256
52
+ verification, Docker startup, `/health`, `/v1/models`, and
53
+ `/v1/chat/completions`.
54
+
55
  This repository contains a single, deployment-ready GGUF optimized and tested
56
  for batch-1 agent workloads on AMD Ryzen AI MAX+ 395 / Radeon 8060S (`gfx1151`,
57
  128 GiB UMA) with the Vulkan backend of llama.cpp b9994. The same artifact was
WINDOWS_ENDPOINT_QUICKSTART.md CHANGED
@@ -34,26 +34,45 @@ weights in host RAM.
34
 
35
  ## 2. Download the full MTP artifact
36
 
37
- Create a model directory and download this repository's GGUF:
 
 
38
 
39
  ```powershell
40
- py -m pip install --upgrade "huggingface_hub[cli]"
41
-
42
  New-Item -ItemType Directory -Force C:\Models\Ornith-MTP | Out-Null
43
 
44
- hf download `
 
 
 
 
 
45
  petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF `
46
  ornith-1.0-35b-MTP-graft-down-Q4_0.gguf `
47
  --local-dir C:\Models\Ornith-MTP
48
  ```
49
 
 
 
 
 
 
 
 
 
 
50
  Expected file:
51
 
52
  ```text
53
  C:\Models\Ornith-MTP\ornith-1.0-35b-MTP-graft-down-Q4_0.gguf
 
54
  SHA-256: 365a7c02dfd320b9696f189d6dc12bd2b0eabb9f8e58ba9fc8cab3af93c0234b
55
  ```
56
 
 
 
 
 
57
  Verify it:
58
 
59
  ```powershell
@@ -227,4 +246,3 @@ docker run --rm `
227
  ```
228
 
229
  It must report llama.cpp version 10066 (`86a9c79f8`) for this exact recipe.
230
-
 
34
 
35
  ## 2. Download the full MTP artifact
36
 
37
+ Create a model directory and an isolated virtual environment for the Hugging
38
+ Face CLI. Do not upgrade `huggingface_hub` in a shared Python installation:
39
+ new CLI releases can require a newer `click` than packages such as `gTTS`.
40
 
41
  ```powershell
 
 
42
  New-Item -ItemType Directory -Force C:\Models\Ornith-MTP | Out-Null
43
 
44
+ $HfVenv = 'C:\Models\Ornith-MTP\.hf-cli'
45
+ py -m venv $HfVenv
46
+ & "$HfVenv\Scripts\python.exe" -m pip install --upgrade pip
47
+ & "$HfVenv\Scripts\python.exe" -m pip install "huggingface_hub[hf_xet]==1.24.0"
48
+
49
+ & "$HfVenv\Scripts\hf.exe" download `
50
  petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF `
51
  ornith-1.0-35b-MTP-graft-down-Q4_0.gguf `
52
  --local-dir C:\Models\Ornith-MTP
53
  ```
54
 
55
+ The download is approximately 19.4 GiB. Xet may spend time scanning chunks
56
+ without continuously printing progress; leave the command running until the
57
+ PowerShell prompt returns. A partially downloaded file is resumed on the next
58
+ identical command.
59
+
60
+ This creates a real standalone copy downloaded from Hugging Face. The recipe
61
+ does not use a hard link, symbolic link, or an already installed LM Studio
62
+ model.
63
+
64
  Expected file:
65
 
66
  ```text
67
  C:\Models\Ornith-MTP\ornith-1.0-35b-MTP-graft-down-Q4_0.gguf
68
+ Size: 20,329,342,112 bytes (18.933 GiB)
69
  SHA-256: 365a7c02dfd320b9696f189d6dc12bd2b0eabb9f8e58ba9fc8cab3af93c0234b
70
  ```
71
 
72
+ This clean download and endpoint flow was re-tested on Windows on 19 July
73
+ 2026, including checksum verification, Docker startup, `/health`, `/v1/models`,
74
+ and `/v1/chat/completions`.
75
+
76
  Verify it:
77
 
78
  ```powershell
 
246
  ```
247
 
248
  It must report llama.cpp version 10066 (`86a9c79f8`) for this exact recipe.