SYCL backend: any speculative type collapses performance (even target prefill drops ~200x) - draft model itself is healthy

#50
by Yoo00ooOO - opened

Environment

  • GPU: Intel Arc A770 16GB (DG2/ACM-G10), driver 32.0.101.8991
  • Stack: oneAPI 2026.1, Level Zero backend (confirmed [level_zero:gpu:0] in banner)
  • OS: Windows 11, Python-free CLI build
  • Source: PrismML fork source snapshot (tarball, build banner reports b0-unknown)
  • Model: Ternary-Bonsai-2-27B PTQ1_0 (qwen35 hybrid arch, 1.75 bpw)
  • Draft: Qwen3.8-4B-Distill Q4_K_M (empero-ai, same vocab, qwen35 hybrid arch)

Summary

With any speculative type enabled (draft-simple, ngram-simple), performance
collapses far below the no-spec baseline β€” and notably the target model's own
prefill collapses too
, which points at a synchronization/pacing problem in the
speculative driver rather than draft quality or GPU offload.

Measurements (same machine, same binaries)

Config pp512 tg128 / generation
Target alone (llama-bench) 215 t/s 15.9 t/s
Draft alone (llama-bench, Qwen3.8-4B Q4_K_M) 1308 t/s 22.1 t/s
spec draft-simple, draft-n-max 16 (llama-cli) 5.9 t/s 0.5–0.9 t/s
spec ngram-simple, draft-n-max 16 (llama-cli, no draft model) 6.4 t/s 9.0 t/s

Notes:

  • The draft model is fully healthy in isolation: offloaded 34/34 layers to GPU,
    1308 t/s prefill. Vocab is compatible (accepted tokens decode correctly).
  • Under draft-simple, verbose timings show draft_n: 6–13, draft_n_accepted: 1–3
    (~20–30% accept rate) and predicted_per_token_ms of 1285–2183 ms β€” vs
    ~63 ms/token for the target alone. The per-step cost scales with draft_n, and
    the draft alone should cost ~50 ms per drafted token at its standalone speed.
  • Under ngram-simple there is no draft model at all, yet target prefill
    still drops from ~1300-class to 6.4 t/s. This is the strongest hint that the
    spec driver enforces a per-step (or per-op) sync on the target queue on SYCL,
    e.g. waiting on every verification batch instead of pipelining.

Reproduce

# baseline (healthy)
llama-bench -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 999 -p 512 -n 128

# collapsed
llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
  -md Qwen3.8-4B-Q4_K_M.gguf -ngl 999 -ngld 999 \
  --spec-type draft-simple --spec-draft-n-max 16 \
  -st -p "say OK" -n 32

# same collapse with no draft model
llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 999 \
  --spec-type ngram-simple --spec-draft-n-max 16 -st -p "say OK" -n 32

Secondary issue: default type inference crashes on plain drafts

Without an explicit --spec-type, a plain (non-sidecar) draft GGUF is routed
into the draft-mtp implementation and aborts:

speculative.cpp:2082: GGML_ASSERT(n_embd == llama_model_n_embd_out(...) &&
  "MTP input row width must match the target h_nextn width") failed

For a draft model with no MTP metadata, falling back to draft-simple (or
failing with a clear "no sidecar found, pass --spec-type" message) would be
friendlier than an assert.

Question

Is SYCL a supported backend for the rewritten speculative driver? If known-broken,
a docs note would save others the debugging session; if not known, happy to
provide verbose logs / tracing output on request.

Sign up or log in to comment