I was hoping for a prune/distill of this model!

#1
by sometimesanotion - opened

Thank you for creating this. I'm very much looking forward to trying it! Will you be publishing an expert index mapping? Have you considered using Unsloth's or Mudler's quantization schemes to make GGUFs of this model?

Tks for your appreciation for my project, those are just demo checkpoints btw, im aiming to compact the model to just 20gb vram while maintaining ~98% of the coding ability

How did u do this? i am trying to do same with glm flash 5.3 such that it fits my 12 gb gpu and 32 gb ddr5

need it only for coding/agentic work, small while retaining frontier level expertise.

would be great if u posted some nice benchmarks for comparision .. thanks

Sign up or log in to comment