Instructions to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF", filename="ornith-1.0-35b-MTP-graft-down-Q4_0.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Use Docker
docker model run hf.co/petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
- LM Studio
- Jan
- vLLM
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
- Ollama
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with Ollama:
ollama run hf.co/petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
- Unsloth Studio
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF to start chatting
- Pi
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with Docker Model Runner:
docker model run hf.co/petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
- Lemonade
How to use petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF:Q4_0
Run and chat with the model
lemonade run user.Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF-Q4_0
List all available models
lemonade list
docs: verify clean Windows download and endpoint recipe
Browse files- README.md +5 -0
- WINDOWS_ENDPOINT_QUICKSTART.md +23 -5
|
@@ -47,6 +47,11 @@ the measured CPU-MoE/GPU-dense MTP profile.
|
|
| 47 |
Endpoint: `http://127.0.0.1:18081/v1`. Do not substitute the no-MTP LM Studio
|
| 48 |
file if speculative acceleration is required.
|
| 49 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 50 |
This repository contains a single, deployment-ready GGUF optimized and tested
|
| 51 |
for batch-1 agent workloads on AMD Ryzen AI MAX+ 395 / Radeon 8060S (`gfx1151`,
|
| 52 |
128 GiB UMA) with the Vulkan backend of llama.cpp b9994. The same artifact was
|
|
|
|
| 47 |
Endpoint: `http://127.0.0.1:18081/v1`. Do not substitute the no-MTP LM Studio
|
| 48 |
file if speculative acceleration is required.
|
| 49 |
|
| 50 |
+
The documented flow was re-tested from a fresh `C:\Models\Ornith-MTP`
|
| 51 |
+
directory using a real Hugging Face download (no hard links), SHA-256
|
| 52 |
+
verification, Docker startup, `/health`, `/v1/models`, and
|
| 53 |
+
`/v1/chat/completions`.
|
| 54 |
+
|
| 55 |
This repository contains a single, deployment-ready GGUF optimized and tested
|
| 56 |
for batch-1 agent workloads on AMD Ryzen AI MAX+ 395 / Radeon 8060S (`gfx1151`,
|
| 57 |
128 GiB UMA) with the Vulkan backend of llama.cpp b9994. The same artifact was
|
|
@@ -34,26 +34,45 @@ weights in host RAM.
|
|
| 34 |
|
| 35 |
## 2. Download the full MTP artifact
|
| 36 |
|
| 37 |
-
Create a model directory and
|
|
|
|
|
|
|
| 38 |
|
| 39 |
```powershell
|
| 40 |
-
py -m pip install --upgrade "huggingface_hub[cli]"
|
| 41 |
-
|
| 42 |
New-Item -ItemType Directory -Force C:\Models\Ornith-MTP | Out-Null
|
| 43 |
|
| 44 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF `
|
| 46 |
ornith-1.0-35b-MTP-graft-down-Q4_0.gguf `
|
| 47 |
--local-dir C:\Models\Ornith-MTP
|
| 48 |
```
|
| 49 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 50 |
Expected file:
|
| 51 |
|
| 52 |
```text
|
| 53 |
C:\Models\Ornith-MTP\ornith-1.0-35b-MTP-graft-down-Q4_0.gguf
|
|
|
|
| 54 |
SHA-256: 365a7c02dfd320b9696f189d6dc12bd2b0eabb9f8e58ba9fc8cab3af93c0234b
|
| 55 |
```
|
| 56 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
Verify it:
|
| 58 |
|
| 59 |
```powershell
|
|
@@ -227,4 +246,3 @@ docker run --rm `
|
|
| 227 |
```
|
| 228 |
|
| 229 |
It must report llama.cpp version 10066 (`86a9c79f8`) for this exact recipe.
|
| 230 |
-
|
|
|
|
| 34 |
|
| 35 |
## 2. Download the full MTP artifact
|
| 36 |
|
| 37 |
+
Create a model directory and an isolated virtual environment for the Hugging
|
| 38 |
+
Face CLI. Do not upgrade `huggingface_hub` in a shared Python installation:
|
| 39 |
+
new CLI releases can require a newer `click` than packages such as `gTTS`.
|
| 40 |
|
| 41 |
```powershell
|
|
|
|
|
|
|
| 42 |
New-Item -ItemType Directory -Force C:\Models\Ornith-MTP | Out-Null
|
| 43 |
|
| 44 |
+
$HfVenv = 'C:\Models\Ornith-MTP\.hf-cli'
|
| 45 |
+
py -m venv $HfVenv
|
| 46 |
+
& "$HfVenv\Scripts\python.exe" -m pip install --upgrade pip
|
| 47 |
+
& "$HfVenv\Scripts\python.exe" -m pip install "huggingface_hub[hf_xet]==1.24.0"
|
| 48 |
+
|
| 49 |
+
& "$HfVenv\Scripts\hf.exe" download `
|
| 50 |
petr567/Ornith-1.0-35B-MTP-Strix-Halo-Hybrid-GGUF `
|
| 51 |
ornith-1.0-35b-MTP-graft-down-Q4_0.gguf `
|
| 52 |
--local-dir C:\Models\Ornith-MTP
|
| 53 |
```
|
| 54 |
|
| 55 |
+
The download is approximately 19.4 GiB. Xet may spend time scanning chunks
|
| 56 |
+
without continuously printing progress; leave the command running until the
|
| 57 |
+
PowerShell prompt returns. A partially downloaded file is resumed on the next
|
| 58 |
+
identical command.
|
| 59 |
+
|
| 60 |
+
This creates a real standalone copy downloaded from Hugging Face. The recipe
|
| 61 |
+
does not use a hard link, symbolic link, or an already installed LM Studio
|
| 62 |
+
model.
|
| 63 |
+
|
| 64 |
Expected file:
|
| 65 |
|
| 66 |
```text
|
| 67 |
C:\Models\Ornith-MTP\ornith-1.0-35b-MTP-graft-down-Q4_0.gguf
|
| 68 |
+
Size: 20,329,342,112 bytes (18.933 GiB)
|
| 69 |
SHA-256: 365a7c02dfd320b9696f189d6dc12bd2b0eabb9f8e58ba9fc8cab3af93c0234b
|
| 70 |
```
|
| 71 |
|
| 72 |
+
This clean download and endpoint flow was re-tested on Windows on 19 July
|
| 73 |
+
2026, including checksum verification, Docker startup, `/health`, `/v1/models`,
|
| 74 |
+
and `/v1/chat/completions`.
|
| 75 |
+
|
| 76 |
Verify it:
|
| 77 |
|
| 78 |
```powershell
|
|
|
|
| 246 |
```
|
| 247 |
|
| 248 |
It must report llama.cpp version 10066 (`86a9c79f8`) for this exact recipe.
|
|
|