New version?

#1
by AI-Joe-git - opened

Hello thanks for sharing this 😁

Is this working now no more missing layer?

Hello thanks for sharing this 😁

Is this working now no more missing layer?

Q8_K_KL is working for me with the latest mtp-clean branch.

Got ~19 t/s boost (64 -> 83) with MTP on my machine for this model.

The current latest mtp-clean branch seems to be broken atm. The latest working commit in the branch that works for me is https://github.com/ggml-org/llama.cpp/commit/e7b4848151377395b1693d326d1cda3fcd61c2d9.

Will there be new version with master llama.cpp now that it's been merged?

Will there be new version with master llama.cpp now that it's been merged?

@zzxc3c I don't think a new quant is needed.

I've just tried Qwen3.6-35B-A3B-heretic.EMB_BF16.Q8_K_XL.gguf,
and it works on llama.cpp b9189 (2026-05-16) πŸ˜€

Sign up or log in to comment