Q2_K_P... yes please

#2
by sa-ai - opened

Your Q4_K_P model is in my top 5 tested recently (past 2 months <=40B), thank you for your efforts.

The original balanced HauhauCS Q2_K_P was showing strong results too, so it carries weight as well... especially since users can run it on a single 16 gig card at 64k+ ctx. MTP would be the cherry on top. This should allow for >30t/s on a 5060 ti.

Thanks again

Me either, for me is the best option for coding, it's very smart.

for anyone else interested... I was able to pull this off (mtp for 27B, any Q from 3.5 or 3.6 Qwen) using 27B_MTP.gguf and convert.py from https://huggingface.co/havenoammo/Qwen3.6-27B-MTP-UD-GGUF
Went from 33 to 56t/s (llama.cpp: --spec-type draft-mtp --spec-draft-n-max 2) on dual 3090s w/nvlink. getting about 50t/s on dual 5060ti's

He has one also for the 35B

Sign up or log in to comment