File size: 7,226 Bytes
8ed94fb
 
 
 
 
 
 
 
 
 
 
 
 
 
ae48755
8ed94fb
 
 
 
 
ae48755
 
 
 
8ed94fb
b090c63
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8ed94fb
 
 
 
 
ae48755
 
 
 
 
8ed94fb
 
 
ae48755
 
 
 
 
 
 
 
 
 
 
 
 
 
8ed94fb
 
 
ae48755
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8ed94fb
 
 
 
 
 
 
ae48755
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8ed94fb
ae48755
8ed94fb
ae48755
 
 
 
8ed94fb
b090c63
 
 
 
 
 
 
8ed94fb
 
 
ae48755
 
 
 
8ed94fb
 
 
ae48755
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
---
license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/ukisai/Swift1.5-Qwen3.8-Flash-Next/blob/0bd4fe22431372cdad1979267d3ab45aa7e6150a/LICENSE
base_model: ukisai/Swift1.5-Qwen3.8-Flash-Next
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
  - gguf
  - mixed-quant
  - qwen4exp
  - qwen3.8-flash-next
  - swift1.5
  - ds4
  - dgx-spark
  - ssd-offload
---

# Swift1.5-Qwen3.8-Flash-Next Mixed-Quant GGUF

> **One smaller Q5 compute backbone with an official FP8 SSD-PLE sidecar.**
> The main precision map matches the published Qwen3.8 and Uncensored
> `MQ-Q5-SSD-PLE-BF16` recipe. All 128 Swift BF16 PLE parts match Qwen,
> allowing the existing official FP8 PLE files to be reused unchanged.

## Performance reference: Qwen base Q5

![Qwen3.8 base Q5 BF16/FP8 PLE throughput reference; not measured on Swift](assets/qwen38-ds4-dfm-rs-2k-64k-throughput.png)

*Reproduced unchanged from the [Qwen model card](https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-SSD-PLE-GGUF/blob/f92da0134498acfefe55aac7bd29e4dcdbbdbfe7/README.md).
Measured on Qwen base Q5, not Swift: one DGX Spark, 2 GiB PLE cache,
16 workers, MTP draft 2; 2,048-token incremental prefill and 128 greedy
tokens per frontier from 2K to 64K. Curves show per-frontier medians;
bands show the observed min–max across three fresh-process runs.*

Swift uses the same graph, tensor shapes, mixed-quant recipe and identical
FP8 PLE files, so similar runtime throughput is expected under matched
conditions. Equal speed has not been measured: routing, generated tokens,
PLE access patterns and MTP acceptance can differ with the main weights.
The BF16 curve is a Qwen comparison; this Swift package ships FP8 PLE only.
[Benchmark protocol, raw data and limits](https://github.com/Baekpica/ds4-dfm-rs/blob/407509a06c366e372499817397918d6690bedcc4/docs/qwen38-ple-fp8.md).

## Support my work

Support model conversion, inference optimization, profiling and open validation:
[Buy Me a Coffee](https://www.buymeacoffee.com/baekpica) · [GitHub Sponsors](https://github.com/sponsors/Baekpica).

Independent Baekpica conversion of
[ukisai/Swift1.5-Qwen3.8-Flash-Next](https://huggingface.co/ukisai/Swift1.5-Qwen3.8-Flash-Next/tree/0bd4fe22431372cdad1979267d3ab45aa7e6150a).
The main weights use the same smaller `MQ-Q5-SSD-PLE-BF16` tensor recipe as
the published Qwen3.8 and Uncensored SSD-PLE models. Conversion ran on CPU
from source BF16 weights, with no imatrix, pruning, expert dropping or layer dropping.

The memory hierarchy follows the existing
[Qwen3.8 mixed-quant SSD-PLE release](https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-SSD-PLE-GGUF):
the 51.2B-parameter PLE table resides on SSD with a bounded runtime page cache,
while the 128.8B-parameter compute backbone uses a mixed-precision GGUF.

## Variant and storage

| Artifact | Storage |
|---|---:|
| Main GGUF, 3 shards / 1,628 tensors | 83,274,984,576 bytes / 77.5559 GiB |
| Main tensor payload | 83,263,928,920 bytes / 77.5456 GiB |
| FP8 PLE, 4 files / 128 logical parts | 51,200,393,216 bytes / 47.6841 GiB |
| FP8 scale | 2 bytes, BF16 `0x3951` |
| Total weight files, including scale | 134,475,377,794 bytes / 125.2400 GiB |

GGUF metadata/alignment, PLE cache, KV state and runtime workspaces are separate costs.

## Quantization targets

| Region | Type |
|---|---|
| Routed gate/up, interior layers 2–45 | Q4_K |
| Routed gate/up, edge layers 0, 1, 46, 47 | Q5_K |
| Routed down, main 512 columns | Q5_K |
| Routed down, 128-column tail | Q5_0 |
| MTP and most always-active matrices | Primarily Q8_0 |
| Non-quantizable matrices, norms and controls | BF16 / F32 / I64 |
| PLE n-gram table on SSD | FP8 E4M3FN |

The exact tensor map is in
[`main-gguf-source-map.json`](reproduction/manifests/swift/ssd-ple-q5/main-gguf-source-map.json).
The main GGUF retains embedded MTP, vision tensors, integer PLE controls and
the source chat template. The standalone tokenizer and processor assets are also included.

## Published artifact layout

```text
MQ-Q5-SSD-PLE-BF16/       Three main GGUF shards, recipe and checksums
PLE-FP8/                 Four FP8 embedding files, scale and manifest
reproduction/            Converter, tensor maps and verification reports
```

## PLE provenance

All 128 Swift BF16 PLE parts match the Qwen BF16 reference. Complete pinned
source-shard SHA-256 proves 124 parts; direct tensor-range SHA-256 proves the
four parts in the two modified mixed shards. The
[`identity report`](reproduction/manifests/swift/verify-ple-identity.json) records both methods.

The official [Qwen FP8 PLE](https://huggingface.co/Qwen/Qwen3.8-Flash-Next-FP8/tree/236dfdf285828023ca3bcd3f37366c58a3469b13)
was copied from the local Qwen model after that complete identity proof.
It is the same FP8 sidecar used by the Baekpica Qwen3.8 and Uncensored releases.
It is not a separate Swift FP8 checkpoint or a new BF16-to-FP8 quantization.

## Download and use with ds4-dfm-rs

This is the dedicated `qwen4exp` SSD-PLE format. Use
[`Baekpica/ds4-dfm-rs`](https://github.com/Baekpica/ds4-dfm-rs), with its FP8 PLE loader;
generic GGUF runtime compatibility is not established.

```bash
swift_repo=Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
swift_root=./Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
hf download "$swift_repo" --local-dir "$swift_root"
(cd "$swift_root" && sha256sum -c SHA256SUMS)
export DS4_QWEN_PLE_DIR="$(realpath "$swift_root/PLE-FP8")"
./ds4 --cuda -m "$swift_root/MQ-Q5-SSD-PLE-BF16/Swift1.5-Qwen3.8-Flash-Next-MQ-Q5-SSD-PLE-BF16-00001-of-00003.gguf" \
  -c 8192 -n 128 -p "Explain binary search."
```

Set `DS4_QWEN_PLE_DIR` explicitly: this package contains FP8 PLE, while the
main packaging name and input precision remain `MQ-Q5-SSD-PLE-BF16`.

## Verification

Source-file checksums, all 1,658 source tensors, the 1,628-tensor Q5 recipe,
PLE controls/layout, complete BF16 PLE identity, main GGUF structure and
independent main/FP8 hashes passed. See [`artifact-manifest.json`](artifact-manifest.json),
[`SHA256SUMS`](SHA256SUMS) and [`reproduction/`](reproduction/README.md).

After conversion, bounded DGX Spark CUDA serving checks passed with 262,144
configured context, two continuous banks, prefill 8,192, MTP draft 2, a 2 GiB
FP8 PLE cache and 32 GiB SSD KV capacity. Text, simultaneous two-bank requests,
prefix/tool continuation reuse and image input passed with zero memory faults.
A full-length 262K prompt and cross-process disk restore were not tested.
Swift throughput and task-quality evaluations have not been run; the graph
above remains a Qwen reference, not a Swift measurement.

## License and attribution

Swift Contribution: [Swift Open License v1.0](LICENSE), Copyright 2026 UkisAI.
Base model: [Qwen Community License 1.0](LICENSE-QWEN), Copyright 2026 Qwen.
Both licenses and the original notices are included. [`NOTICE`](NOTICE) records
the conversion changes. Quantizer code is MIT licensed; see [`ds4/LICENSE`](ds4/LICENSE).

Thanks to UkisAI for Swift1.5, Qwen for the base model and FP8 PLE, GGML/llama.cpp
for the quantization foundations, and the DwarfStar/ds4 contributors.
This conversion is not an official UkisAI or Qwen release.