pipenetwork's picture
Add 5bit tier to the measured quantization comparison
15f8b33 verified
|
Raw
History Blame
7.29 kB
metadata
language:
  - en
  - zh
license: apache-2.0
library_name: mlx
pipeline_tag: image-text-to-text
tags:
  - mlx
  - apple-silicon
  - qwen3.5
  - fine tune
  - heretic
  - uncensored
  - abliterated
  - merge
  - thinking
  - reasoning
  - creative
  - writing
  - fiction
  - roleplaying
  - vision
base_model: nightmedia/Qwen3.5-9B-DS9-USS-Defiant

Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-MLX-4bit

MLX conversion of Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic for Apple silicon — 4-bit, the smallest tier

Uncensored ("Heretic'd") multi-stage merge of Qwen3.5-9B fine tunes by nightmedia and DavidAU, with a compacted-but-stronger thinking block. Vision is included and works out of the box — no separate mmproj download.

Provenance

This was converted from nightmedia/Qwen3.5-9B-DS9-USS-Defiant (bfloat16 safetensors), which is the exact same weight set that DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF packages as GGUF.

We verified the identity numerically rather than assuming it:

Check Result
lm_head.weight (BF16 in both) vs GGUF output.weight bit-exact, 0 mismatches across 262,144 values
RMSNorm weights (input_layernorm, q_norm, post_attention_layernorm, final norm) equal source + 1.0 under llama.cpp's shift convention (12542/12544 elements exact; 2 differ by 1 ULP from the float32 addition)
embed_tokens vs GGUF Q8_0 token_embd cosine 0.999956 — consistent with a plain Q8_0 round-trip

So these MLX quants are made from the original bf16 weights, not by dequantizing a GGUF. There is no GGUF round-trip loss.

Credit for the model itself goes to nightmedia and DavidAU; this repo only does the MLX conversion.

Architecture

Qwen3.5-9B is a hybrid multimodal model:

  • 32 layers, 3:1 ratio of gated-delta linear attention to full attention (full_attention_interval: 4)
  • Gated output attention (attn_output_gate), head_dim 256, 16 Q heads / 4 KV heads
  • Interleaved mRoPE with partial_rotary_factor 0.25, rope_theta 1e7
  • 262,144 native context
  • 27-layer vision tower (patch 16, spatial merge 2), 248,320 vocab

The source also ships MTP (multi-token-prediction) weights. MLX does not use them, so they are dropped during conversion — this costs no quality, only the speculative-decoding speedup that the GGUF "MTP" variants offer.

Usage

Vision + text with mlx-vlm:

pip install mlx-vlm
python -m mlx_vlm.generate \
  --model pipenetwork/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-MLX-4bit \
  --image your_image.jpg \
  --prompt "Describe this image in detail." \
  --max-tokens 512

Text-only with mlx-lm (loads the same repo, ignores the vision tower):

pip install mlx-lm
mlx_lm.generate \
  --model pipenetwork/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-MLX-4bit \
  --prompt "Write the opening paragraph of a noir story set on a space station." \
  --max-tokens 512

Python:

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("pipenetwork/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-MLX-4bit")
config = model.config

prompt = apply_chat_template(processor, config, "Describe this image.", num_images=1)
print(generate(model, processor, prompt, ["your_image.jpg"], max_tokens=512, verbose=False))

Quantizations

Repo Bits Size
...-MLX-4bit 4 6.0 GB
...-MLX-6bit 6 7.7 GB
...-MLX-5bit 5 6.6 GB
...-MLX-8bit 8 9.7 GB
...-MLX-bf16 16 18.8 GB

Group size 64, affine mode. The vision tower is left unquantized (mlx-vlm's default for multimodal projector/patch-embed modules), so the size delta between tiers comes from the language model.

Measured quantization quality

Against the bf16 reference these were converted from — 65,536 tokens of wikitext-2 test at 1024 context, identical token ids through mlx-lm, on an M3 Ultra. KL is measured against bf16's own output distribution, so lower means closer to the original model.

model lang. weights ppl Δppl KL(bf16‖quant) top-1 vs bf16 decode
bf16 reference 17.91 GB 8.1273 39.2 t/s
8bit 9.51 GB 8.1277 +0.00% 0.00124 98.24% 67.9 t/s
6bit 7.28 GB 8.1426 +0.19% 0.00523 96.20% 80.9 t/s
5bit 6.16 GB 8.2012 +0.91% 0.01845 93.24% 92.5 t/s
4bit (this repo) 5.04 GB 8.5816 +5.59% 0.07330 87.06% 110.0 t/s
nightmedia mxfp4 4.76 GB 8.9443 +10.05% 0.11328 82.34% 114.7 t/s

Reading it:

  • 8bit is effectively free — bf16 perplexity to four decimals at 47% of the footprint and 1.7x the decode speed.
  • 5bit is the best quality-per-GB: under 1% perplexity for 6.2 GB.
  • The cliff is 5bit → 4bit, where KL jumps 4x and perplexity goes +0.91% → +5.59%. If 4bit feels lossy, 5bit is the tier to move to, not 6bit.
  • At the 4-bit tier this affine group-64 quant loses about half the perplexity MXFP4 does, costing 4.5 vs 4.25 bits/weight (~6% more storage, ~4% slower decode).

All tiers are quantizations of the same weights, so this isolates the quantization scheme. Reproduce with bench.py.

Sampling## Sampling

DavidAU's notes for this model, which carry over:

  • Temperature 1.0 or below works best; higher temps degrade coherence
  • Repetition penalty 1.0 (off) — raising it hurts this model
  • The thinking block is compacted; give it room with a generous max-tokens

Benchmarks

From the source model card (source weights, non-MLX quant tiers), for reference:

          arc/c  arc/e boolq hswag obkqa piqa  wino
bf16      0.649, 0.832, 0.895, 0.713, 0.482, 0.783, 0.699
mxfp8     0.647, 0.836, 0.895, 0.706, 0.460, 0.784, 0.695
mxfp4     0.640, 0.824, 0.886, 0.703, 0.468, 0.780, 0.691

Qwen3.5-9B-Instruct (base, non-heretic)
mxfp8     0.571, 0.719, 0.895, 0.683, 0.426, 0.770, 0.671

These are not measurements of these MLX repos — treat them as characterising the weights, not this quantization.

Conversion tooling

Scripts, benchmark and the provenance proof: github.com/PipeNetwork/defiant-fable-mlx

License

Apache 2.0, inherited from Qwen3.5-9B.

This model has had its safety post-training removed and will follow instructions without refusal. You are responsible for how you use it.