Incredible project questions about total download size,coding quality,benchmarks

#4
by Abdulazizv - opened

Awesome project , and
Am i correct that You are trying to make model size smaller while preserving agentic coding quality it will be close to the original model right ?
And also will the final model be able to run as fast as other 35b moe models.
do you plan to do benchmarks
I am really hyped for this project .
Fully supporting you !!

yep, but im not so sure if im going to do a full benchmark comparison since i only have around 76gb vram. (will consider about that later on )

following this thread

Hi ,
saw your discussion where you asked and I would lean toward option B, because the single gpu version would make this much more accesible to users with limited vram.

However, i also think the bf16 dora work is also really valuable.if the failures are like mostly edge cases rather than capability issues it means tuning can could make it bf16 checkpoint a much stronger foundation.
I think ideal path is to keep the bf16 checkpoint as the main reference model and make int4 int8 variabts from the improved version.
It would be also interesting to see int4 vs int8 results on the same 100 tasks
This is already a pretty fascinating thing.

good idea, imma try that soon

Hey Jab, just wondering β€” are you still planning to release the ~19.7GB INT4/smaller version?

I saw that the current model file is around 40GB, while the README mentions ~19.7GB. I'd really love to test the smaller version locally on my limited-VRAM setup.

No pressure, just curious if you're still planning to make it. Really cool project btw!

ye im still working on it, it is just that i have been quite busy recently so i cant release it yet. i still need to finetune the model first then i can quantize it down to Q4. i will try release it asap for you guys

Sign up or log in to comment