Anyone tried? How do you think?

#13
by evilperson068 - opened

Is it a good model?
I'd love to hear your thoughts, thanks!

Just tried the gguf 17gb k-quant with llama.cpp. Feels quite stupid compared to Qwen3.6-27B-Fable-Fus-711-UnHeretic-NM-DAU-NEO-MAX-NEO-MTP-Q4_K_M. You can also see this in the reasoning, where it seems to forget and re-iterate on things. In the pi harness, at one point it also stashed away the changes I wanted it to fix. That's where I pulled the plug.

On the good side, when using the dflash drafter, it does more than 60 tok/sec on my 3090. (A bit faster than with Qwen 3.6 w/ MTP.) Also, vision (with the additional mmproj model) seems to work really well.

Just tried the gguf 17gb k-quant with llama.cpp. Feels quite stupid compared to Qwen3.6-27B-Fable-Fus-711-UnHeretic-NM-DAU-NEO-MAX-NEO-MTP-Q4_K_M. You can also see this in the reasoning, where it seems to forget and re-iterate on things. In the pi harness, at one point it also stashed away the changes I wanted it to fix. That's where I pulled the plug.

On the good side, when using the dflash drafter, it does more than 60 tok/sec on my 3090. (A bit faster than with Qwen 3.6 w/ MTP.) Also, vision (with the additional mmproj model) seems to work really well.

Any sign of bench-maxxing?

I have no idea on how to test that.

Sign up or log in to comment