Please do not quant the token embeddings to q4_0

#10
by Dampfinchen - opened

There is a reason why Google sets them to q6_k even for QAT. Lower quality embeddings will not show up in benchmarks but only in real world use cases with a long context size. Errors are tiny at first but accumulate the longer the context grows. So what you are doing is pretty risky for just a few 100 mb's less. I really would not risk it.

Sign up or log in to comment