It is fighting with the harness, chat template and blaming the human operator

#35
by gbuzhf - opened

Nah, doesn't work. I'm on recommended sampling ... o_O

Hermes
image

Opencode2
image

What's wrong with it?
Totally lost.

Yeah they quantized every single bit of the model to 1.72BPW and that breaks the model fully. Fitting on 8GB was probably their priority but squeezing every sensitive bit during the process makes this model useless. They only needed to add 1.44GiB of file to preserve emebddings and lm heads at q6_k and total just 3.3GiB more to keep every single sensitive part at q6_k including embedding, lm heads, Gated deltanets and Gated global attentions. It just loops for tens of thousands of token when I use Xhigh reasoning. Bonsai models have been very disappointing. Qwen3.5 9b is better for 8gb cards.

Edit: 35b a3b indeed is faster and better. Use qwen3.6 35b a3b instead of this mess.

Thanks for joining, Qwen3.5 9b is better for 8gb cards. -> 35B-A3B is pretty much near-perfect for 8GB cards and works fast enough (27-40 t/s).
And yes, this 27b try is really useless. Deleted already.

Thanks for joining, Qwen3.5 9b is better for 8gb cards. -> 35B-A3B is pretty much near-perfect for 8GB cards and works fast enough (27-40 t/s).
And yes, this 27b try is really useless. Deleted already.

Yeah. I forgot 35b a3b was an option. Honestly fitting model entirely on 8GB Vram is very difficult. They should have targeted 12GB or 16GB cards because at 3 bit post-training quantization starts to become unstable and QAT can recover almost all of it if done correctly.

Really,
Tell DeepSeek follow these instructions: https://huggingface.co/peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF/discussions/18#6aaee757b2001a3d2ba2bdee put TheTom's fork here C:\tools\turboquant\.
for this https://huggingface.co/gbuzhf/Ornith-1.5-35B-A3B-Huihui-Sangreal-MTP-ICE-GGUF/blob/main/Ornith-1.5-35B-A3B-Huihui-Sangreal-MTP-23G-ICE.gguf model.
with this vision https://huggingface.co/AesSedai/Qwen3.6-35B-A3B-GGUF/blob/main/mmproj-Qwen3.6-35B-A3B-Q8_0.gguf. --n-cpu-moe 38 is good for RAM preservation.
if you've 8GB VRAM & 32GB RAM, it'll just work well, speed and capability of the model will amaze you, be little tolerant with prefill (400-600) that's all.

I've tried both 9B and 35B; 9B is nowhere near it.

21G-ICE for more RAM. :)

Speed/Image reading:
image

Yeah they quantized every single bit of the model to 1.72BPW and that breaks the model fully. Fitting on 8GB was probably their priority but squeezing every sensitive bit during the process makes this model useless. They only needed to add 1.44GiB of file to preserve emebddings and lm heads at q6_k and total just 3.3GiB more to keep every single sensitive part at q6_k including embedding, lm heads, Gated deltanets and Gated global attentions. It just loops for tens of thousands of token when I use Xhigh reasoning. Bonsai models have been very disappointing. Qwen3.5 9b is better for 8gb cards.

Edit: 35b a3b indeed is faster and better. Use qwen3.6 35b a3b instead of this mess.

on Xhigh this specific qwen model goes through everything and loops for a long time, that happens with the f16 model too, and a q6_k weight packed into this wouldn't make a ternary quant but whose counting numbers here, it's not like we're doing math, we're all just passing opinions around.

Yeah they quantized every single bit of the model to 1.72BPW and that breaks the model fully. Fitting on 8GB was probably their priority but squeezing every sensitive bit during the process makes this model useless. They only needed to add 1.44GiB of file to preserve emebddings and lm heads at q6_k and total just 3.3GiB more to keep every single sensitive part at q6_k including embedding, lm heads, Gated deltanets and Gated global attentions. It just loops for tens of thousands of token when I use Xhigh reasoning. Bonsai models have been very disappointing. Qwen3.5 9b is better for 8gb cards.

Edit: 35b a3b indeed is faster and better. Use qwen3.6 35b a3b instead of this mess.

on Xhigh this specific qwen model goes through everything and loops for a long time, that happens with the f16 model too, and a q6_k weight packed into this wouldn't make a ternary quant but whose counting numbers here, it's not like we're doing math, we're all just passing opinions around.

That's not true. This ternary model endlessly thought for 110k tokens until it hit output limit(just to make a flappy bird). I ran q5 of original version and it didn't think for this long(21k thinking). Both with recommended thinking sampling parameters with repeat panelty 1. Even with repeat panelty set to 1.1 to prevent looping, the result was not much comparable to result from q5 version. With medium reasoning it doesn't loop as much but it degrades output even more.

Yeah they quantized every single bit of the model to 1.72BPW and that breaks the model fully. Fitting on 8GB was probably their priority but squeezing every sensitive bit during the process makes this model useless. They only needed to add 1.44GiB of file to preserve emebddings and lm heads at q6_k and total just 3.3GiB more to keep every single sensitive part at q6_k including embedding, lm heads, Gated deltanets and Gated global attentions. It just loops for tens of thousands of token when I use Xhigh reasoning. Bonsai models have been very disappointing. Qwen3.5 9b is better for 8gb cards.

Edit: 35b a3b indeed is faster and better. Use qwen3.6 35b a3b instead of this mess.

on Xhigh this specific qwen model goes through everything and loops for a long time, that happens with the f16 model too, and a q6_k weight packed into this wouldn't make a ternary quant but whose counting numbers here, it's not like we're doing math, we're all just passing opinions around.

That's not true. This ternary model endlessly thought for 110k tokens until it hit output limit(just to make a flappy bird). I ran q5 of original version and it didn't think for this long(21k thinking). Both with recommended thinking sampling parameters with repeat panelty 1. Even with repeat panelty set to 1.1 to prevent looping, the result was not much comparable to result from q5 version. With medium reasoning it doesn't loop as much but it degrades output even more.

you are comparing 5 float points to ternary, of course f16 gonna be even better than 5 float points, and much better than ternary. a quant is always a give and take, with ternary you give more time so you can be able to even run it with a gaming chip. assuming you used the XL version of a q5 quant that's ~4x the size of the ternary model. meaning you need to have 4x the GPUs to just say hello and hit a OOM before it even finishes thinking.
you're not passing unbiased data, you're passing opinions that are truly worthless for a person with a low end chip.

Yeah they quantized every single bit of the model to 1.72BPW and that breaks the model fully. Fitting on 8GB was probably their priority but squeezing every sensitive bit during the process makes this model useless. They only needed to add 1.44GiB of file to preserve emebddings and lm heads at q6_k and total just 3.3GiB more to keep every single sensitive part at q6_k including embedding, lm heads, Gated deltanets and Gated global attentions. It just loops for tens of thousands of token when I use Xhigh reasoning. Bonsai models have been very disappointing. Qwen3.5 9b is better for 8gb cards.

Edit: 35b a3b indeed is faster and better. Use qwen3.6 35b a3b instead of this mess.

on Xhigh this specific qwen model goes through everything and loops for a long time, that happens with the f16 model too, and a q6_k weight packed into this wouldn't make a ternary quant but whose counting numbers here, it's not like we're doing math, we're all just passing opinions around.

That's not true. This ternary model endlessly thought for 110k tokens until it hit output limit(just to make a flappy bird). I ran q5 of original version and it didn't think for this long(21k thinking). Both with recommended thinking sampling parameters with repeat panelty 1. Even with repeat panelty set to 1.1 to prevent looping, the result was not much comparable to result from q5 version. With medium reasoning it doesn't loop as much but it degrades output even more.

you are comparing 5 float points to ternary, of course f16 gonna be even better than 5 float points, and much better than ternary. a quant is always a give and take, with ternary you give more time so you can be able to even run it with a gaming chip. assuming you used the XL version of a q5 quant that's ~4x the size of the ternary model. meaning you need to have 4x the GPUs to just say hello and hit a OOM before it even finishes thinking.
you're not passing unbiased data, you're passing opinions that are truly worthless for a person with a low end chip.

I know what quantization is. Just saying it is totally incorrect to compare useless 110k of thinking vs useful 21k of thinking. That is what you are implying. Btw my opinion is totally valid for mid end pc. I have mid end laptop and I was excited but disappointed by the abysmal performance of the model with real harnesses. At least hoped it to get iq3_m performance but it is on par with low bit q2. The 98.2% figure is so misleading.

Yeah they quantized every single bit of the model to 1.72BPW and that breaks the model fully. Fitting on 8GB was probably their priority but squeezing every sensitive bit during the process makes this model useless. They only needed to add 1.44GiB of file to preserve emebddings and lm heads at q6_k and total just 3.3GiB more to keep every single sensitive part at q6_k including embedding, lm heads, Gated deltanets and Gated global attentions. It just loops for tens of thousands of token when I use Xhigh reasoning. Bonsai models have been very disappointing. Qwen3.5 9b is better for 8gb cards.

Edit: 35b a3b indeed is faster and better. Use qwen3.6 35b a3b instead of this mess.

on Xhigh this specific qwen model goes through everything and loops for a long time, that happens with the f16 model too, and a q6_k weight packed into this wouldn't make a ternary quant but whose counting numbers here, it's not like we're doing math, we're all just passing opinions around.

That's not true. This ternary model endlessly thought for 110k tokens until it hit output limit(just to make a flappy bird). I ran q5 of original version and it didn't think for this long(21k thinking). Both with recommended thinking sampling parameters with repeat panelty 1. Even with repeat panelty set to 1.1 to prevent looping, the result was not much comparable to result from q5 version. With medium reasoning it doesn't loop as much but it degrades output even more.

you are comparing 5 float points to ternary, of course f16 gonna be even better than 5 float points, and much better than ternary. a quant is always a give and take, with ternary you give more time so you can be able to even run it with a gaming chip. assuming you used the XL version of a q5 quant that's ~4x the size of the ternary model. meaning you need to have 4x the GPUs to just say hello and hit a OOM before it even finishes thinking.
you're not passing unbiased data, you're passing opinions that are truly worthless for a person with a low end chip.

I know what quantization is. Just saying it is totally incorrect to compare useless 110k of thinking vs useful 21k of thinking. That is what you are implying. Btw my opinion is totally valid for mid end pc. I have mid end laptop and I was excited but disappointed by the abysmal performance of the model with real harnesses. At least hoped it to get iq3_m performance but it is on par with low bit q2. The 98.2% figure is so misleading.

you're telling me that your +20Gbs VRAM GPU (likely a 5090) is mid end. if your reporting about your own device is anything like your opinions then we definitely have a good sample about the quality and validity of your opinions. I'm muting this conversation.

Yeah they quantized every single bit of the model to 1.72BPW and that breaks the model fully. Fitting on 8GB was probably their priority but squeezing every sensitive bit during the process makes this model useless. They only needed to add 1.44GiB of file to preserve emebddings and lm heads at q6_k and total just 3.3GiB more to keep every single sensitive part at q6_k including embedding, lm heads, Gated deltanets and Gated global attentions. It just loops for tens of thousands of token when I use Xhigh reasoning. Bonsai models have been very disappointing. Qwen3.5 9b is better for 8gb cards.

Edit: 35b a3b indeed is faster and better. Use qwen3.6 35b a3b instead of this mess.

on Xhigh this specific qwen model goes through everything and loops for a long time, that happens with the f16 model too, and a q6_k weight packed into this wouldn't make a ternary quant but whose counting numbers here, it's not like we're doing math, we're all just passing opinions around.

That's not true. This ternary model endlessly thought for 110k tokens until it hit output limit(just to make a flappy bird). I ran q5 of original version and it didn't think for this long(21k thinking). Both with recommended thinking sampling parameters with repeat panelty 1. Even with repeat panelty set to 1.1 to prevent looping, the result was not much comparable to result from q5 version. With medium reasoning it doesn't loop as much but it degrades output even more.

you are comparing 5 float points to ternary, of course f16 gonna be even better than 5 float points, and much better than ternary. a quant is always a give and take, with ternary you give more time so you can be able to even run it with a gaming chip. assuming you used the XL version of a q5 quant that's ~4x the size of the ternary model. meaning you need to have 4x the GPUs to just say hello and hit a OOM before it even finishes thinking.
you're not passing unbiased data, you're passing opinions that are truly worthless for a person with a low end chip.

I know what quantization is. Just saying it is totally incorrect to compare useless 110k of thinking vs useful 21k of thinking. That is what you are implying. Btw my opinion is totally valid for mid end pc. I have mid end laptop and I was excited but disappointed by the abysmal performance of the model with real harnesses. At least hoped it to get iq3_m performance but it is on par with low bit q2. The 98.2% figure is so misleading.

you're telling me that your +20Gbs VRAM GPU (likely a 5090) is mid end. if your reporting about your own device is anything like your opinions then we definitely have a good sample about the quality and validity of your opinions. I'm muting this conversation.

Please stop this joke. Obviously I have both PC and Laptop. Stop assuming everything in your favor. I have 3090 desktop and laptop is RX 6800M 12GB.

Sign up or log in to comment