Community MLX conversions with native MTP

#4
by vsnake87 - opened

Thanks to huihui-ai for creating and releasing the original model. These are independent community conversions of that work, maintained by KostkaIT.

I've prepared two independent MLX conversions of this model for local inference on Apple Silicon:

Both versions preserve the model's native MTP weights and are designed to work with oMLX without requiring a separate drafter model. I prefer this self-contained setup because it is simpler to use and maintain.

In my own testing, the oQ4e build retained almost all of the quality I was getting from Q8, while using less memory. Results will depend on the hardware, runtime, and settings, so feedback from other users is welcome.

These are independent community conversions maintained by KostkaIT. They are not official releases from Huihui or Qwen.

Sign up or log in to comment