4bpw feedback

#1
by TheodoreH - opened

Hi. Since there is no posibilities to run models bigger than 15 gb - stuck with that one, and having issues. It loops quite frequently, but for 4bpw it should be much better(using qwen 35b with 3.5 bits and it had never looped at all, but GLM loops almost every third answer and if message is going long(none code one) it loops too). tried different settings, but lower temps or pens - fix(somewhat) hallucinating and loosing grip of logic in text, but it is not perfect, or not at least fine, bearable for now. it is shame, because GLM seems to write in very different manner than Gemma or Qwen, and it is the only reason it makes sense to try it more. Still thanks for upload, because there are none really similar quants for this model. Usually use Apex quants or Cerebellum, but there are none fore unfiltered GLM that can be used normally. All availible (less than 15gb) quants or IQ3M - or 3 KM - is far worse and dont work normally at all. From DavidAU it runs much slower for some reason(but model is same).

Sign up or log in to comment