Q8?

#2
by MB7977 - opened

Hi,

Thank you for these. Any plans to reinstate the Q8? I have the one you deleted but presume it needs updating? I have tried the 5KM but this model is extremely sensitive to quantization and I found the Q8 made a quality difference, so prefer to run as close to full precision as possible. No worries if you don’t plan to, just need to know if I should stop checking. :)

Sorry i already deleted the safetensors and bf16 cause i needed room on my drive and hf. I thought Aes uploaded one but usually 5km mixed is almost as good. Might want to check to see if he will upload his.

Once it gets mainlined hopefully soon, bartowski will have one

Tbh though, if you can run the Q5 it's still smart enough in a terminal agent session to quant the Q8 for you. You also can quant straight from safetensors to Q8. I pretty much orchestrated all these via deepseek-flash in pi

Oh wait lcpp doesnt have tool calling for this yet. Anyway most models can do it for you 🤗🤗🤗

Sure, thanks. Yeah I know how to quant manually. Just trying to avoid downloading the full weights and going through the process.

MB7977 changed discussion status to closed

Sign up or log in to comment