This is the conversion of z-lab/Qwen3.6-27B-DFlash to be used with llama.cpp folowing the method in the first post https://github.com/ggml-org/llama.cpp/pull/22105 Offered as a convenience.

DFlash is a second model to load as a drafter.

The command : As of June 29, I tested it against ROCm, Vulkan Fails. I expect CUDA to work

llama-server \
-m /models/Qwen3.6-27B-GGUF-4.256bpw-imatrix.gguf \
--spec-draft-n-max 8 \
--spec-type draft-dflash --spec-draft-model /models/qwen3.6-27b-dflash-IQ4_XS.gguf \
-a "qwen-27b" \
--host 127.0.0.1 \
--port 17102 \
--device ROCm0 \
-c 40000 \
-ngl 40 \
Downloads last month
509
GGUF
Model size
2B params
Architecture
dflash
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jojohai/Qwen3.6-27B-DFlash-GGUF

Quantized
(18)
this model