Why is file size different compared to mradermacher's quants?

#3
by BigBeavis - opened

First time I notice something like this. Same model, and yet bartowski quants are all bigger in filesize than mradermachers? I thought this might be due to how you set up the configuration for some tensors, but even Q8 is different in filesize! Whut? Q6_K is 2gb different in size. Even K_S quants all different in size. Sorry to bother you, just want to understand what's actually going on.

First time I notice something like this. Same model, and yet bartowski quants are all bigger in filesize than mradermachers? I thought this might be due to how you set up the configuration for some tensors, but even Q8 is different in filesize! Whut? Q6_K is 2gb different in size. Even K_S quants all different in size. Sorry to bother you, just want to understand what's actually going on.

I assume Bartowski left some more layers at Q8_0 because he noticed degraded performance when quantised to Q4_0.

That doesn't explain why Q8 different in size as well

Sign up or log in to comment