Q8 Version low t/s

#6
by FearL0rd - opened

I have the usloth q8 version running at 40 t/s but this model is doing 8 t/s for the same prompt using llama.cpp

do you have an idea whats going on?

Sign up or log in to comment