Instructions to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Use Docker
docker model run hf.co/SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SEBK4C/Ornith-1.0-35B-MTP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SEBK4C/Ornith-1.0-35B-MTP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
- Ollama
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with Ollama:
ollama run hf.co/SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with Docker Model Runner:
docker model run hf.co/SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
- Lemonade
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ornith-1.0-35B-MTP-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use SEBK4C/Ornith-1.0-35B-MTP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SEBK4C/Ornith-1.0-35B-MTP-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Ornith-1.0-35B MTP GGUF (Q4_K_M / Q6_K / Q8_0)
- TL;DR — which file to download
- The finding: a quant-sensitivity ladder
- MTP / speculative decoding notes
- Serving (what we actually run)
- Fixed chat template (
serving/ornith-35b-fixed.jinja) - Known model-family bugs (all quants, reproducible)
- Files
- Ornith-Serve: one-file tuned serving (Linux x86_64 + NVIDIA)
- Provenance & attribution
- Honest limitations
- TL;DR — which file to download
Ornith-1.0-35B MTP GGUF (Q4_K_M / Q6_K / Q8_0)
Three MTP-grafted quantizations of deepreinforce-ai/Ornith-1.0-35B — each shipped with a measured agentic verdict, not a perplexity chart. All three were benchmarked head-to-head on identical machine-scored task batteries (multi-turn agentic, long-horizon reasoning, code-with-hidden-tests; full harness + journals: SEBK4C/ornith-35b-agentic-eval).
TL;DR — which file to download
| file | bpw | decode (2×4090, MTP) | multi-turn agentic | hard battery | verdict |
|---|---|---|---|---|---|
| Q6_K-MTP | 6.56 | 242 t/s solo | 6/6 | 7/10 | Use this one. |
| Q8_0-MTP | 8.50 | 189 t/s (−22%) | 6/6-equiv | 6/10 | No measured advantage over Q6 anywhere (~40 cells, 2 runs). Thinks +13-40% more tokens without converging better. |
| Q4_K_M-MTP | 4.89 | 200 t/s (+11% vs Q8-matched) | 4/6 | 5/10 | Published as a warning artifact — read below before using. |
The finding: a quant-sensitivity ladder
On this RL-post-trained agentic model, quantization damage does not hit capabilities uniformly. Measured order of fragility:
- Agentic persistence breaks first (Q4) — Q4_K_M aces every single-shot reasoning cell (58k-token needle chains, 17-hop code traces, 99-step recurrences) but abandons multi-turn tasks: 2-3 turns in, required output files never written, trajectory just ends. The RL-trained "diligence" is the most quant-fragile thing in the model — and no perplexity metric shows it.
- Format precision gets noisy next (Q8-vs-Q6 trajectory level) — exact verbatim-output discipline (keep a required prefix, exact line counts) flips with seed at both Q6 and Q8; 3-seed study: Q6 5/6 vs Q8 3/6 on the discordant cells.
- Reasoning is robust all the way down — exact-answer math/logic/retrieval held at every quant we tested.
Practical rule: don't requant an agentic RL model below Q6 without a multi-turn diligence eval. Single-shot benchmarks will lie to you.
Why is Q8 not better than Q6?
Across two independent 10-task hard-battery runs plus a 4-cell brutal battery: Q8_0 never won a single cell against Q6_K, ran 22% slower, and consumed +13-40% more thinking tokens per battery. A control run (official Q8_0, no MTP, different llama.cpp build) reproduced the Q8-MTP failures cell-for-cell — excluding serving machinery. Hypothesis (not proven): Q8's rollouts deliberate longer without converging better on this always-reasoning model; part of the observed gap is seed-level trajectory divergence (quantified in the dataset repo's seed-variance journals). What is proven: there is no upside to paying Q8's cost.
MTP / speculative decoding notes
- All three files carry the 20-tensor MTP head (
blk.40.*inclnextn.*). Serve with--spec-type draft-mtp --spec-draft-n-max 3on a llama.cpp with MTP support (merged 2026-05, PR #22673). - MTP survives requantization: draft acceptance measured identical (47%) on Q4 and Q6 with matched prompts. The head is not the fragile part.
- MTP inverts under batch: measured crossover at
2 concurrent streams (with MTP: 205/220/256/286 t/s agg at 1/2/4/8 streams; without: 171/250/ 349/451; at 16: 321 vs 514). Enable MTP only for workloads that are mostly 1-2 concurrent; disable for fleets. (serving/*.argsencodes this.) -np 1is NOT required on current builds (stale early-notes claim; verified to np16+).
Serving (what we actually run)
See serving/ for complete flag sets. Summary (2×24GB GPUs):
| profile | np × ctx/slot | MTP | measured |
|---|---|---|---|
fastest.args |
1 × 262144 | on | 242 t/s solo, full native context |
fast.args |
3 × 131072 | on | 208 solo / 248 agg@3 |
parallel.args |
16 × 65536 | off | 514 t/s agg@16 (613@32 works too) |
- Sampling (model card):
--temp 0.6 --top-p 0.95 --top-k 20. - KV:
q4_0measured clean at 58k-token retrieval depth; the hybrid-attention architecture makes KV ~9.4GB per 1M tokens at q4_0 — context is cheap. --reasoning-budget 10240recommended for work profiles: the model overthinks; the budget force-closes<think>so an answer always emits (converts silent token-cap deaths into honest attempts).- Instant no-think mode per request:
"chat_template_kwargs": {"enable_thinking": false}— ~2× cheaper on mechanical multi-step tool work, near-parity quality; keep thinking ON for combinatorial planning.
Fixed chat template (serving/ornith-35b-fixed.jinja)
The template embedded in upstream GGUFs hard-raises (HTTP 500 through llama-server) on three real-world client behaviors. This repo ships a fixed template, byte-identical to upstream on all happy paths (offline render-diff proven), lenient on:
developerrole → rendered as a system block (OpenAI SDK / Codex clients send this; upstream raisesUnexpected message role).- Mid-conversation
systemmessages → rendered in place (upstream raises). - Tool-continuation histories with no plain user query → rendered (upstream
raises
No user query found in messages; the guard'slast_query_indexis dead code in the 35B template).
Serve with --chat-template-file ornith-35b-fixed.jinja --jinja --reasoning-format deepseek.
Known model-family bugs (all quants, reproducible)
-2**2evaluates to 4: the model consistently binds unary minus tighter than**(8/8 attempts across Q4/Q6/Q8, with and without self-testing). Self-verification does not fix it — the model writes its tests from the same false belief. Harness rule: supply golden test vectors in prompts; never prose-only specs.- Unaided full-context aggregation (~700 records) exceeds reliable capability
at any quant/reasoning depth — but a shell workspace solves it in ~1.4k
tokens via self-scaffolded
grep|awk(the model was RL-trained to generate scaffolds; harness it shell-first).
Files
| file | bytes | sha256 |
|---|---|---|
| Ornith-1.0-35B-Q6_K-MTP.gguf | 29,207,678,080 | 449288fa43409a791058ba134b83f4dd10174ef2e6dd63044047fe07960d6834 |
| ornith-1.0-35b-Q8_0-MTP.gguf | 37,596,664,960 | 837cd5a1b40b4e7fb437911564f36fcf4885540a453421cf1450b35d4b51f627 |
| ornith-1.0-35b-Q4_K_M-MTP.gguf | 21,712,409,728 | 54eb0fcc203abf974984adfc0188d96624954481ac684ec533df3faa551f43ab |
Ornith-Serve: one-file tuned serving (Linux x86_64 + NVIDIA)
Ornith-Serve-x86_64-cuda.run (29GB, self-extracting) bundles a patched
llama.cpp server + CUDA runtime libs + the Q6_K-MTP weights + the fixed
template + all three measured profiles behind one launcher:
./Ornith-Serve-x86_64-cuda.run # extract once
./ornith-serve/ornith-serve.sh fastest # or: fast | parallel
What the bundled build has that stock llama.cpp doesn't: MTP self-speculative
decoding for this model, the fixed chat template, measured profiles, and an
experimental --spec-max-active batch-occupancy gate for speculative decoding
(default off; drafts only while ≤N streams generate — the one-server-all-modes
design is in progress, see Honest limitations). Requires the NVIDIA driver;
~24GB VRAM (fastest), 2×24GB recommended for parallel.
sha256: 6a30cfba93c715ebe82d5de7a3dd486f08e4399eee353fa9ef1d106e3cdb2b73
A true cross-OS llamafile is planned once llamafile's llama.cpp sync includes MTP (its current engine predates it).
Provenance & attribution
- Base model: deepreinforce-ai/Ornith-1.0-35B (MIT) — agentic-coding RL post-train of Qwen3.5-35B MoE. Official GGUFs ship without the MTP tensors.
- MTP graft method: skinnyctax/Ornith-1.0-35B-Q6_K-Frankenstein-MTP-GGUF —
20 head tensors appended, metadata patched (
block_count+1,nextn_predict_layers=1). No weights retrained. - Q6_K-MTP file: re-hosted from wang-yang/Ornith-1.0-35B-MTP-GGUF (MIT), unchanged.
- Q8_0-MTP: same graft procedure applied to the official Q8_0 GGUF.
- Q4_K_M-MTP:
llama-quantize --allow-requantizefrom the Q8_0-MTP file (deterministic; requant noise is part of the published finding).
Honest limitations
- Benchmarks are single-seed unless marked; the 3-seed variance study covers the Q6-vs-Q8 discordant cells only. Cell counts are small (6-24 per battery); we publish direction + full journals, not significance claims.
- All numbers from one host (2×RTX 4090, llama.cpp-mtp build 4fc4ec5 era).
--spec-max-active(in the Ornith-Serve build) gates drafting correctly and preserves solo speed, but spec-armed state currently carries a fixed per-iteration cost inside llama_decode that caps gated batch throughput (~223 vs 514 t/s @16 streams). Until that's profiled out, use theparallelprofile (spec off) for fleets — the gate is for burst protection on small-slot profiles.- The eval tasks are freshly authored (contamination-resistant) but ours; the full harness is published for scrutiny and reruns.
- Downloads last month
- 254
4-bit
6-bit
8-bit
Model tree for SEBK4C/Ornith-1.0-35B-MTP-GGUF
Base model
ornith-ai/Ornith-1.0-35B