For DGX Spark (GB10) owners: a vLLM alternative to GGUF with the native DSpark speculative head - ~10 tok/s, 256K ctx, tool calling

#22
by wiklif - opened

Great to see 0731 GGUFs here - for anyone running them on a single DGX Spark class box (GB10, 128GB unified memory), I want to share a complementary path I got working and published: serving this model from the original checkpoint through a vLLM-based stack, with things the GGUF route currently cannot use.

https://github.com/lrozewicz/vLLM-Moet-GB10

It builds on vLLM-Moet (2-bit expert planes + an FP4 tier that recovers quality where 2-bit loses it), adapted for unified memory and sm_121. What you get over a plain Q2 GGUF setup:

  • the model's built-in DSpark speculative head (the extra 20B in this checkpoint that llama.cpp does not use): ~9.8 tok/s decode on a single GB10 (k=2; heads-up - the model card's k=7 is actually slower on this hardware, 5.1 tok/s)
  • 256K context served (KV pool ~660K tokens thanks to MLA + the compression ladder)
  • an OpenAI-compatible server with working DSML tool calling (--tool-call-parser=deepseek_v4) and a per-request thinking toggle with split reasoning - drop-in backend for coding agents
  • first boot ~35 min (one-time quantization), later boots ~10 min via a persistent cache

The README documents the unified-memory pitfalls that make this non-obvious on GB10 (pinned host staging blowing past 121 GiB, misleading memory metrics, etc.), each with the measurement behind it.

Not a replacement for GGUF - llama.cpp remains the simpler route - but if you have a Spark and want the speculative head, long context and agent tooling, this is reproducible end to end. Feedback welcome.

I don't know, your performance doesn't look that good. I have the NVidia Thor system, basically the little brother of the Spark (only half the CUDA cores).

$ LD_LIBRARY_PATH=./ ./llama-bench --model "/space/models/unsloth:DeepSeek-V4-Flash-0731-UD-IQ3_S.gguf" -fa on -lm dio -p 2048 -n 512
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 125771 MiB):
  Device 0: NVIDIA Thor, compute capability 11.0, VMM: yes, VRAM: 125771 MiB
load_backend: loaded CUDA backend from /space/llama/libggml-cuda.so
load_backend: loaded RPC backend from /space/llama/libggml-rpc.so
load_backend: loaded CPU backend from /space/llama/libggml-cpu.so
| model                          |       size |     params | backend    | ngl |  fa |         lm |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| deepseek4 ?B IQ3_S - 3.4375 bpw | 108.09 GiB |   284.33 B | CUDA       |  -1 |   1 |        dio |          pp2048 |        154.90 ± 0.34 |
| deepseek4 ?B IQ3_S - 3.4375 bpw | 108.09 GiB |   284.33 B | CUDA       |  -1 |   1 |        dio |           tg512 |         14.49 ± 0.17 |

build: ddd4ec142 (10217)

Yeah, it is the 3bit version. I can easily run a context window of half a million tokens. And because of the increasing amount of active parameters with a growing context window I get about 7 token/s, if the window is filled to 350k tokens. In the beginning the speed is about 14.5 token/s and goes down to about 10 token/s when 100k tokens are filled. Your Spark should actually pull about 1.5 times the numbers. The Deepseek V4 Flash model has 13b active parameters. The Thor and Spark have a memory bandwidth of 273 GB/s (or 253 GiB/s). LLMs are just simple forward passes of all active parameters, aka, pure memory streaming. If all the active parameters are purely 8bit (Q8_0), the maximum possible performance is: 273b/13b = 21 token/s (well, plus the parameters from the context window)
Though, Deepseek V4 is about 95% fp4 or 4bit, means the theoretical maximum should be about 42 token/s.
42 token/s * 13b parameters doing basically a MAC operation (multiply-add, in case of a GPU = 2 FLOPS or OPS in case of INTS) for every single parameter results in 1092 GFLOPS/GOPS or ~1.1 TOPS. So computing power shouldn't be a problem. Both machines have a serious memory bandwidth limit. But you use the mostly 2bit quantized version on a much stronger hardware, your performance looks odd.

Sign up or log in to comment