Quantizing larger models?

#1
by mindplay - opened

Are you able to quantize larger models, say 27B or 36B models, to 1 or 1.5 bits?

Since you're only releasing your own models, it's difficult to say how much of your benchmark advantage is due to your model and training data/approach, versus how much is due to your quantization algorithms, isn't it?

Having an 8B model I can run on my phone is kind of cool.

Having a 27B or maybe even 36B model I can run on my 4060 Ti with 16 GB (at actually useful speeds!) would be next level. πŸ˜„

Sign up or log in to comment