GGUF
conversational
How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf John1604/Qwen3-Next-80B-A3B-Instruct-gguf:
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "John1604/Qwen3-Next-80B-A3B-Instruct-gguf:" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Qwen3 Next Instruct gguf

Make sure you have enough memory/GPU

Use the model in ollama

First download and install ollama.

https://ollama.com/download

Note: the official ollama models do not have Qwen3-Next support yet. You need do the following.

Command

in windows command line (or mac os, linux), or in terminal in ubuntu, type:

ollama run hf.co/John1604/Qwen3-Next-80B-A3B-Instruct-gguf:q3_k_m

(q3_k_m is the model quant type, q3_k_s, q4_k_m, ..., can also be used)

C:\Users\developer>ollama run hf.co/John1604/Qwen3-Next-80B-A3B-Instruct-gguf:q3_k_m
pulling manifest
...
writing manifest
success

>>> Send a message (/? for help) 

After you run command: ollama run hf.co/John1604/Qwen3-Next-80B-A3B-Instruct-gguf:q3_k_m, it will appear in ollama UI - you may select this model hf.co/John1604/Qwen3-Next-80B-A3B-Instruct-gguf:q3_k_m from the model list, and run it the same way as other ollama supported models.

Use the model in LM Studio

download and install LM Studio

https://lmstudio.ai/

Discover models

In the LM Studio, click "Discover" icon. "Mission Control" popup window will be displayed.

In the "Mission Control" search bar, type "John1604/Qwen3-Next-80B-A3B-Instruct-gguf" and check "GGUF", the model should be found.

Download a quantized model.

Load the quantized model.

Ask questions.

quantized models

Type Bits Quality Description
Q2_K 2-bit 🟥 Low Minimal footprint; only for tests
Q3_K_S 3-bit 🟧 Low “Small” variant (less accurate)
Q3_K_M 3-bit 🟧 Low–Med “Medium” variant
Q4_K_S 4-bit 🟨 Med Small, faster, slightly less quality
Q4_K_M 4-bit 🟩 Med–High “Medium” — best 4-bit balance
Q5_K_S 5-bit 🟩 High Slightly smaller than Q5_K_M
Q5_K_M 5-bit 🟩🟩 High Excellent general-purpose quant
Q6_K 6-bit 🟩🟩🟩 Very High Almost FP16 quality, larger size
Q8_0 8-bit 🟩🟩🟩🟩 Near-lossless baseline
Downloads last month
130
GGUF
Model size
80B params
Architecture
qwen3next
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for John1604/Qwen3-Next-80B-A3B-Instruct-gguf

Quantized
(73)
this model