Ternary-based 35B-A3B I made, using this distill.

#8
by SkyIsNotGreen - opened

Thanks for the distill. I quantized Qwen3.8-35B-A3B-Distill down to 2.61 bpw and wanted to show you the result.

Scion-35B-A3B: the expert banks in a ternary container (2-bit codes plus one fp16 group scale per 128 weights), everything else Q8_0, plus small trained rank-512 corrections (attention output, MoE block output and router deltas), trained by output-KD against the BF16 teacher with the deployed quantizer in the loop. No imatrix, no calibration corpus, no full-model QAT. The corrections are the only trained part, and the run cost about $7 of rented H100 time.

  • 11.34 GB / 2.61 bpw, single file (corrections embedded, no --lora), 6.3x smaller than the BF16 reference.
  • Task retention (400 tasks each): HellaSwag 79.00 and Winogrande 76.25 against BF16's 81.25 and 76.00, inside the Β±2% noise band; best PPL of the 2-bit class (8.354 versus IQ2_M's 8.413).
  • The gap: full-vocabulary KLD against BF16 is still 2-bit-class (0.269 mean versus Q4_K_M's 0.031). Tail-aware training is my next step.

Apache-2.0, inherited, with attribution to empero-ai and Qwen (and the community BF16 GGUF conversion). It needs my llama.cpp fork (PQ2_0 plus embedded adapters; stock llama.cpp cannot load the tensor types). I built it to serve a local assistant stack on a single 20 GB card.

Card and weights: https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B
Build-scripts, write-up and failed routes: https://github.com/sky-is-green/scion

If anyone wants to try it or poke holes in the numbers, the model repo's discussions are open, and I am happy to share any part of the recipe in more detail.

Do you plan to use SignRoundV2 and/or CAT-Q? https://arxiv.org/html/2512.04746v2 https://arxiv.org/html/2606.26650v1

I suspect this Ternary training harness might be useful:
https://huggingface.co/penkia/TernaryQuench-Qwen3.8-27B-GGUF
They did something similar with 27B, but I suspect if SignRoundV2 is integrated the results would improve.

That said, did you use these, and if not what else did you use?

I tell you out of my own self-interest, I'd like the 35B-A3B to use without having to do the work πŸ˜„

Sign up or log in to comment