Instructions to use unsloth/DeepSeek-V4-Flash-0731-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- Ollama
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Ollama:
ollama run hf.co/unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
- Unsloth Studio
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/DeepSeek-V4-Flash-0731-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/DeepSeek-V4-Flash-0731-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for unsloth/DeepSeek-V4-Flash-0731-GGUF to start chatting
- Pi
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Docker Model Runner:
docker model run hf.co/unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
- Lemonade
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.DeepSeek-V4-Flash-0731-GGUF-UD-Q4_K_XL
List all available models
lemonade list
For DGX Spark (GB10) owners: a vLLM alternative to GGUF with the native DSpark speculative head - ~10 tok/s, 256K ctx, tool calling
Great to see 0731 GGUFs here - for anyone running them on a single DGX Spark class box (GB10, 128GB unified memory), I want to share a complementary path I got working and published: serving this model from the original checkpoint through a vLLM-based stack, with things the GGUF route currently cannot use.
https://github.com/lrozewicz/vLLM-Moet-GB10
It builds on vLLM-Moet (2-bit expert planes + an FP4 tier that recovers quality where 2-bit loses it), adapted for unified memory and sm_121. What you get over a plain Q2 GGUF setup:
- the model's built-in DSpark speculative head (the extra 20B in this checkpoint that llama.cpp does not use): ~9.8 tok/s decode on a single GB10 (k=2; heads-up - the model card's k=7 is actually slower on this hardware, 5.1 tok/s)
- 256K context served (KV pool ~660K tokens thanks to MLA + the compression ladder)
- an OpenAI-compatible server with working DSML tool calling (
--tool-call-parser=deepseek_v4) and a per-request thinking toggle with split reasoning - drop-in backend for coding agents - first boot ~35 min (one-time quantization), later boots ~10 min via a persistent cache
The README documents the unified-memory pitfalls that make this non-obvious on GB10 (pinned host staging blowing past 121 GiB, misleading memory metrics, etc.), each with the measurement behind it.
Not a replacement for GGUF - llama.cpp remains the simpler route - but if you have a Spark and want the speculative head, long context and agent tooling, this is reproducible end to end. Feedback welcome.
I don't know, your performance doesn't look that good. I have the NVidia Thor system, basically the little brother of the Spark (only half the CUDA cores).
$ LD_LIBRARY_PATH=./ ./llama-bench --model "/space/models/unsloth:DeepSeek-V4-Flash-0731-UD-IQ3_S.gguf" -fa on -lm dio -p 2048 -n 512
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 125771 MiB):
Device 0: NVIDIA Thor, compute capability 11.0, VMM: yes, VRAM: 125771 MiB
load_backend: loaded CUDA backend from /space/llama/libggml-cuda.so
load_backend: loaded RPC backend from /space/llama/libggml-rpc.so
load_backend: loaded CPU backend from /space/llama/libggml-cpu.so
| model | size | params | backend | ngl | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| deepseek4 ?B IQ3_S - 3.4375 bpw | 108.09 GiB | 284.33 B | CUDA | -1 | 1 | dio | pp2048 | 154.90 ± 0.34 |
| deepseek4 ?B IQ3_S - 3.4375 bpw | 108.09 GiB | 284.33 B | CUDA | -1 | 1 | dio | tg512 | 14.49 ± 0.17 |
build: ddd4ec142 (10217)
Yeah, it is the 3bit version. I can easily run a context window of half a million tokens. And because of the increasing amount of active parameters with a growing context window I get about 7 token/s, if the window is filled to 350k tokens. In the beginning the speed is about 14.5 token/s and goes down to about 10 token/s when 100k tokens are filled. Your Spark should actually pull about 1.5 times the numbers. The Deepseek V4 Flash model has 13b active parameters. The Thor and Spark have a memory bandwidth of 273 GB/s (or 253 GiB/s). LLMs are just simple forward passes of all active parameters, aka, pure memory streaming. If all the active parameters are purely 8bit (Q8_0), the maximum possible performance is: 273b/13b = 21 token/s (well, plus the parameters from the context window)
Though, Deepseek V4 is about 95% fp4 or 4bit, means the theoretical maximum should be about 42 token/s.
42 token/s * 13b parameters doing basically a MAC operation (multiply-add, in case of a GPU = 2 FLOPS or OPS in case of INTS) for every single parameter results in 1092 GFLOPS/GOPS or ~1.1 TOPS. So computing power shouldn't be a problem. Both machines have a serious memory bandwidth limit. But you use the mostly 2bit quantized version on a much stronger hardware, your performance looks odd.