--- license: other license_name: swift-open-license-1.0 license_link: https://huggingface.co/ukisai/Swift1.5-Qwen3.8-Flash-Next/blob/0bd4fe22431372cdad1979267d3ab45aa7e6150a/LICENSE base_model: ukisai/Swift1.5-Qwen3.8-Flash-Next base_model_relation: quantized pipeline_tag: image-text-to-text tags: - gguf - mixed-quant - qwen4exp - qwen3.8-flash-next - swift1.5 - ds4 - dgx-spark - ssd-offload --- # Swift1.5-Qwen3.8-Flash-Next Mixed-Quant GGUF > **One smaller Q5 compute backbone with an official FP8 SSD-PLE sidecar.** > The main precision map matches the published Qwen3.8 and Uncensored > `MQ-Q5-SSD-PLE-BF16` recipe. All 128 Swift BF16 PLE parts match Qwen, > allowing the existing official FP8 PLE files to be reused unchanged. ## Performance reference: Qwen base Q5 ![Qwen3.8 base Q5 BF16/FP8 PLE throughput reference; not measured on Swift](assets/qwen38-ds4-dfm-rs-2k-64k-throughput.png) *Reproduced unchanged from the [Qwen model card](https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-SSD-PLE-GGUF/blob/f92da0134498acfefe55aac7bd29e4dcdbbdbfe7/README.md). Measured on Qwen base Q5, not Swift: one DGX Spark, 2 GiB PLE cache, 16 workers, MTP draft 2; 2,048-token incremental prefill and 128 greedy tokens per frontier from 2K to 64K. Curves show per-frontier medians; bands show the observed min–max across three fresh-process runs.* Swift uses the same graph, tensor shapes, mixed-quant recipe and identical FP8 PLE files, so similar runtime throughput is expected under matched conditions. Equal speed has not been measured: routing, generated tokens, PLE access patterns and MTP acceptance can differ with the main weights. The BF16 curve is a Qwen comparison; this Swift package ships FP8 PLE only. [Benchmark protocol, raw data and limits](https://github.com/Baekpica/ds4-dfm-rs/blob/407509a06c366e372499817397918d6690bedcc4/docs/qwen38-ple-fp8.md). ## Support my work Support model conversion, inference optimization, profiling and open validation: [Buy Me a Coffee](https://www.buymeacoffee.com/baekpica) · [GitHub Sponsors](https://github.com/sponsors/Baekpica). Independent Baekpica conversion of [ukisai/Swift1.5-Qwen3.8-Flash-Next](https://huggingface.co/ukisai/Swift1.5-Qwen3.8-Flash-Next/tree/0bd4fe22431372cdad1979267d3ab45aa7e6150a). The main weights use the same smaller `MQ-Q5-SSD-PLE-BF16` tensor recipe as the published Qwen3.8 and Uncensored SSD-PLE models. Conversion ran on CPU from source BF16 weights, with no imatrix, pruning, expert dropping or layer dropping. The memory hierarchy follows the existing [Qwen3.8 mixed-quant SSD-PLE release](https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-SSD-PLE-GGUF): the 51.2B-parameter PLE table resides on SSD with a bounded runtime page cache, while the 128.8B-parameter compute backbone uses a mixed-precision GGUF. ## Variant and storage | Artifact | Storage | |---|---:| | Main GGUF, 3 shards / 1,628 tensors | 83,274,984,576 bytes / 77.5559 GiB | | Main tensor payload | 83,263,928,920 bytes / 77.5456 GiB | | FP8 PLE, 4 files / 128 logical parts | 51,200,393,216 bytes / 47.6841 GiB | | FP8 scale | 2 bytes, BF16 `0x3951` | | Total weight files, including scale | 134,475,377,794 bytes / 125.2400 GiB | GGUF metadata/alignment, PLE cache, KV state and runtime workspaces are separate costs. ## Quantization targets | Region | Type | |---|---| | Routed gate/up, interior layers 2–45 | Q4_K | | Routed gate/up, edge layers 0, 1, 46, 47 | Q5_K | | Routed down, main 512 columns | Q5_K | | Routed down, 128-column tail | Q5_0 | | MTP and most always-active matrices | Primarily Q8_0 | | Non-quantizable matrices, norms and controls | BF16 / F32 / I64 | | PLE n-gram table on SSD | FP8 E4M3FN | The exact tensor map is in [`main-gguf-source-map.json`](reproduction/manifests/swift/ssd-ple-q5/main-gguf-source-map.json). The main GGUF retains embedded MTP, vision tensors, integer PLE controls and the source chat template. The standalone tokenizer and processor assets are also included. ## Published artifact layout ```text MQ-Q5-SSD-PLE-BF16/ Three main GGUF shards, recipe and checksums PLE-FP8/ Four FP8 embedding files, scale and manifest reproduction/ Converter, tensor maps and verification reports ``` ## PLE provenance All 128 Swift BF16 PLE parts match the Qwen BF16 reference. Complete pinned source-shard SHA-256 proves 124 parts; direct tensor-range SHA-256 proves the four parts in the two modified mixed shards. The [`identity report`](reproduction/manifests/swift/verify-ple-identity.json) records both methods. The official [Qwen FP8 PLE](https://huggingface.co/Qwen/Qwen3.8-Flash-Next-FP8/tree/236dfdf285828023ca3bcd3f37366c58a3469b13) was copied from the local Qwen model after that complete identity proof. It is the same FP8 sidecar used by the Baekpica Qwen3.8 and Uncensored releases. It is not a separate Swift FP8 checkpoint or a new BF16-to-FP8 quantization. ## Download and use with ds4-dfm-rs This is the dedicated `qwen4exp` SSD-PLE format. Use [`Baekpica/ds4-dfm-rs`](https://github.com/Baekpica/ds4-dfm-rs), with its FP8 PLE loader; generic GGUF runtime compatibility is not established. ```bash swift_repo=Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF swift_root=./Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF hf download "$swift_repo" --local-dir "$swift_root" (cd "$swift_root" && sha256sum -c SHA256SUMS) export DS4_QWEN_PLE_DIR="$(realpath "$swift_root/PLE-FP8")" ./ds4 --cuda -m "$swift_root/MQ-Q5-SSD-PLE-BF16/Swift1.5-Qwen3.8-Flash-Next-MQ-Q5-SSD-PLE-BF16-00001-of-00003.gguf" \ -c 8192 -n 128 -p "Explain binary search." ``` Set `DS4_QWEN_PLE_DIR` explicitly: this package contains FP8 PLE, while the main packaging name and input precision remain `MQ-Q5-SSD-PLE-BF16`. ## Verification Source-file checksums, all 1,658 source tensors, the 1,628-tensor Q5 recipe, PLE controls/layout, complete BF16 PLE identity, main GGUF structure and independent main/FP8 hashes passed. See [`artifact-manifest.json`](artifact-manifest.json), [`SHA256SUMS`](SHA256SUMS) and [`reproduction/`](reproduction/README.md). After conversion, bounded DGX Spark CUDA serving checks passed with 262,144 configured context, two continuous banks, prefill 8,192, MTP draft 2, a 2 GiB FP8 PLE cache and 32 GiB SSD KV capacity. Text, simultaneous two-bank requests, prefix/tool continuation reuse and image input passed with zero memory faults. A full-length 262K prompt and cross-process disk restore were not tested. Swift throughput and task-quality evaluations have not been run; the graph above remains a Qwen reference, not a Swift measurement. ## License and attribution Swift Contribution: [Swift Open License v1.0](LICENSE), Copyright 2026 UkisAI. Base model: [Qwen Community License 1.0](LICENSE-QWEN), Copyright 2026 Qwen. Both licenses and the original notices are included. [`NOTICE`](NOTICE) records the conversion changes. Quantizer code is MIT licensed; see [`ds4/LICENSE`](ds4/LICENSE). Thanks to UkisAI for Swift1.5, Qwen for the base model and FP8 PLE, GGML/llama.cpp for the quantization foundations, and the DwarfStar/ds4 contributors. This conversion is not an official UkisAI or Qwen release.