Solved "xhigh" loop and include MTP. Test on 4080 12G VRAM laptop, reach ~60 t/s

#45
by zhijin123 - opened

Lastly, tesing PQ2_0 with "xhigh", i hited thinking loop (maybe very long thinking).
Now, i increased a bit penalty "--repeat-penalty 1.05" with xhigh, the problem seems solved.
Also, i tested MTP. It worked and increased 1x% speed on 4080 12G VRAM laptop to ~60 t/s.
Below is my PrismML-Eng/llama.cpp arg. (it took ~11.4G VRAM) :
.\llama-server -m D:\AI_model\prismML\Ternary-Bonsai-2-27B\Ternary-Bonsai-2-27B-PQ2_0.gguf -md D:\AI_model\prismML\Ternary-Bonsai-2-27B\mtp-Qwen3.8-27B-Q4_0.gguf -ngl all -fa on -np 1 -c 81920 -ctk q8_0 -ctv q8_0 --fit off --temp 0.6 --top-p 0.95 --top-k 20 --repeat-penalty 1.05 --reasoning-effort xhigh --spec-type draft-mtp --spec-draft-n-max 2 --jinja --port 8964

Include the latest evalscope tool_bench testing result (It completed 1 more test and the score is 0.4231)
image

Sign up or log in to comment