Ok hear me out.

#1
by Miltos22 - opened

This is smart, but turning 27b into an moe model would likely be smarter. ram + vram moe can often run better than pure vram dense with less quality loss than 16b. Id attempt it myself but im on vacation

I been thinking about this, I've managed to get a decent 14b but yeah as you said there is a limitation to compression. Busy balancing funding but i am planning on trying to get an MOE variant out. Ironically its more difficult than the by layer compression. Definitely in the works! If you'd like to test the best of the compressed its logic65/Qwen3.8-Whittle-tri-14.7B but still needs some fine-tuning and still uploading .

logic65 changed discussion status to closed
logic65 changed discussion status to open

Okay you won me over. I'm going to see if I can squeeze out something tonight.

Will vouch for this, this is all I can have with my really modest hardware. A finished Qwen3.8 27B MoE would be delightful. Great work so far.

Its a work in progress and currently repairing with KD distill on a rented A100 unfortunately it wont be a 3b active more like 17b but looking into getting that number down as far as possible.

Sign up or log in to comment