The repository contains huihui-ai/Huihui-Qwen3.8-27B-abliterated converted to the AMD Quark MXFP4 format using the AMD specified 28 calibration samples @ 512 sequences.
This model has had the drafter converted to FP8 using https://codeberg.org/ggz14/radiance-vllm-mxfp4/src/branch/main/fp8_mtp.py for the ggz14 radiance-vllm-mxfp4 image (need to compile local and run via run_mxfp4_074.sh). https://codeberg.org/ggz14/radiance-vllm-mxfp4/
The process was driven on a system with 128GB DDR4, 5900x, and 2x R9700 gpus. The last 3 layers had to be offloaded to cpu but it's not yet determined if that affected quality. PPL was slightly higher than the AMD reference model. My testing and any PPL hit the original BF16 model took from uncensoring are suspected factors. Initial testing feels ok with good results.
Reference model: https://huggingface.co/amd/Qwen3.8-27B-Quark-AWQ-MXFP4
License Apache 2.0, inherited from Qwen/Qwen3.8-27B.
- Downloads last month
- 119
Model tree for gearwave00001/Huihui-Qwen3.8-27B-Quark-AWQ-MXFP4-MtpFp8
Base model
Qwen/Qwen3.8-27B