Can we also have FP16 for M1 max and M2 max family? Thank you so much.

#1
by mastervivi - opened

Can we also have FP16 for M1 max and M2 max family? Thank you so much. We really appreciate a lot ❣️

I've run a conversion of this quant to fp16 using a script derived from clone_mlx_model_fp16.py, and uploaded it here

I've run a conversion of this quant to fp16 using a script derived from clone_mlx_model_fp16.py, and uploaded it here

I downloaded this to try it out on my M1 Ultra, but I got a substantial drop in performance, from ~18.5 tok/s down to ~9.7 tok/s, so I cannot recommend using monroewilliam's conversion

Sign up or log in to comment