I wonder how the model would have performed without quantizing embedding and other sensitive parts down to 1.72bpw.

#17
by tigerjjw53 - opened

The embedding quantization breaks multilingual ability pretty hard. It is practically unusable for languages other than english due to token corruption and language mixing. I wonder how the model would have performed if embedding were left at q6_k(0.97GiB) and lm output heads(also 0.97GiB). That only adds 1.44GiB(subtracted 1.72bpw size) of data to the model while preserving vocabulary. + possibly all gated deltanet layers(1.51GiB at q6_k) and gated global attention layers(1.02GiB at q6_k). Total just 3.3GiB(subtracting 1.72bpw size) more to the model while preserving most sentive parts. SwiGLU FFN takes the most space of 13.07GiB at q6_k and like just about 3.45GiB when 1.72bit quantization is applied. The model is currently very brittle and I would love to see next generation sacrifice a bit of space for actually production ready models.

Edit: The multilingual ability seems to be pretty fine when I tighten the top k and min p a bit. Still though, it is noticeably less fluent than base model. Makes up words.

tigerjjw53 changed discussion status to closed
tigerjjw53 changed discussion status to open

sounds goodd, but maybe they were targeting a specific VRAM budget. At this size +1GB or +3GB means +30% or +50%, which is the difference between fitting on an 8GB card or not

sounds goodd, but maybe they were targeting a specific VRAM budget. At this size +1GB or +3GB means +30% or +50%, which is the difference between fitting on an 8GB card or not

It is currently too unstable that running qwen3.5 9b is better for stability. They should have targetted 8GB vram with offload since unusable fast model is worse than usable slower model.

Sign up or log in to comment