Hi,everyone,I wanna know 24GB can do it?
#1
by link921 - opened
Pls,thanks....I don't know
24GB VRAM? Yes, if you use --ssd-streaming flags to stream it from the hard drive. Or, if you have 70GB+ system RAM for the rest of the experts not in VRAM, then definitely yes. Either way, it is possible, but SSD streaming is not very fast, just FYI.
I have 64GB DDR5 RAM + 24GB VRAM totaling 88GB. Leaving some RAM to the system only fits the IQ2XXS 86.7 GB model itself without any context. My question is whether I can perform further quantization of some other weights in the network that can reduce around 3.5GB that I would replace with context. If so, what are the candidate tensor classes to cut down from Q8 to Q4 or from F16 to Q8 sacrificing some hopefully minimal accuracy on the way.