avlp12/GLM-5.3-Flash-Alis-MLX-8bit
Text Generation • Updated • 1.04k • 4
Rewritten 2026-08-30: router/mHC/KDA aux kept unquantized. 4bit=QUASAR-init (KL -8.5% vs RTN). ~29-33 tok/s on M3 Ultra. Receipts in each repo.
Note Standalone native MTP (nextn) drafter, bf16 — 1.12x bit-identical self-speculative decoding via mlx-vlm PR #2044