Qwopus3.8-27B-Flash-V2 · attention8-bf16recurrence

Experimental MLX conversion of Jackrong/Qwopus3.8-27B-Flash-V2, pinned to 13f92e09a46fa364f8de1edb85684d57bda01126. Separate V2 release; no abliteration. Unqualified: known Python-formatting failures remain.

Mixed affine quantization: 234 sensitive modules at 8-bit/group 64, 168 modules at 4-bit/group 32, and 96 recurrent input projections in BF16. Weight files total 22.186 GiB; runtime memory is higher. The package includes 333 same-parent BF16 vision tensors and a 15-tensor native BF16 MTP sidecar. Tokenizer, chat template and generation configuration are preserved; flat image-processor metadata mirrors the source settings. Exact per-module precision is recorded in config.json; file hashes are in SHA256SUMS.

Runtime

Built with MLX 0.32.2 / MLX-LM 0.31.3; packaged for MTPLX 2.11.2 on Apple Silicon. Use MTPLX for the combined text/vision/native-MTP artifact; ordinary MLX-LM text generation does not enable vision or MTP.

hf download Shiftedx/qwopus3.8-27b-flash-v2-attention8-bf16recurrence-vision-mtplx --local-dir model
mtplx inspect --model model --json
mtplx serve --model model --backend-id qwen3_next --generation-mode ar --reasoning-mode on --reasoning-effort xhigh --temperature 0.3 --top-p 0.95 --top-k 20

For native MTP, use --generation-mode mtp --depth 3; D3 was exercised on Attention8 only. Start with AR when evaluating another variant. Pass enable_thinking=true and reasoning_effort="xhigh" explicitly in chat-template/API controls. The preserved source generation config defaults to temperature 1.0; override it to 0.3 for coding. No maximum-context or cross-runtime qualification is claimed.

Evidence and limitations

Loaded and exercised with text, images, tools and native MTP. Local behavioral gates failed: Python indentation, strict output formatting and an agent task. Full publication qualification is incomplete.

One controlled coding prompt also failed indentation in the untouched BF16 parent, loaded through standard MLX-LM with the same seed and recommended sampling settings. Raw token decoding confirmed the defect. This shows that our quantization is not required to trigger that failure; it does not establish a universal source-model failure rate or exclude MLX-specific behavior. The upstream card reports improvement, not guaranteed elimination. No benchmark leaderboard or speed claim is made here.

Other V2 variants:

Downloads last month
21
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shiftedx/qwopus3.8-27b-flash-v2-attention8-bf16recurrence-vision-mtplx

Base model

Qwen/Qwen3.8-27B
Quantized
(10)
this model