Baekpica's picture
Add Qwen throughput reference and Swift serving checks
b090c63 verified
|
Raw History Blame Contribute Delete
7.23 kB
metadata
license: other
license_name: swift-open-license-1.0
license_link: >-
  https://huggingface.co/ukisai/Swift1.5-Qwen3.8-Flash-Next/blob/0bd4fe22431372cdad1979267d3ab45aa7e6150a/LICENSE
base_model: ukisai/Swift1.5-Qwen3.8-Flash-Next
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
  - gguf
  - mixed-quant
  - qwen4exp
  - qwen3.8-flash-next
  - swift1.5
  - ds4
  - dgx-spark
  - ssd-offload

Swift1.5-Qwen3.8-Flash-Next Mixed-Quant GGUF

One smaller Q5 compute backbone with an official FP8 SSD-PLE sidecar. The main precision map matches the published Qwen3.8 and Uncensored MQ-Q5-SSD-PLE-BF16 recipe. All 128 Swift BF16 PLE parts match Qwen, allowing the existing official FP8 PLE files to be reused unchanged.

Performance reference: Qwen base Q5

Qwen3.8 base Q5 BF16/FP8 PLE throughput reference; not measured on Swift

Reproduced unchanged from the Qwen model card. Measured on Qwen base Q5, not Swift: one DGX Spark, 2 GiB PLE cache, 16 workers, MTP draft 2; 2,048-token incremental prefill and 128 greedy tokens per frontier from 2K to 64K. Curves show per-frontier medians; bands show the observed min–max across three fresh-process runs.

Swift uses the same graph, tensor shapes, mixed-quant recipe and identical FP8 PLE files, so similar runtime throughput is expected under matched conditions. Equal speed has not been measured: routing, generated tokens, PLE access patterns and MTP acceptance can differ with the main weights. The BF16 curve is a Qwen comparison; this Swift package ships FP8 PLE only. Benchmark protocol, raw data and limits.

Support my work

Support model conversion, inference optimization, profiling and open validation: Buy Me a Coffee · GitHub Sponsors.

Independent Baekpica conversion of ukisai/Swift1.5-Qwen3.8-Flash-Next. The main weights use the same smaller MQ-Q5-SSD-PLE-BF16 tensor recipe as the published Qwen3.8 and Uncensored SSD-PLE models. Conversion ran on CPU from source BF16 weights, with no imatrix, pruning, expert dropping or layer dropping.

The memory hierarchy follows the existing Qwen3.8 mixed-quant SSD-PLE release: the 51.2B-parameter PLE table resides on SSD with a bounded runtime page cache, while the 128.8B-parameter compute backbone uses a mixed-precision GGUF.

Variant and storage

Artifact Storage
Main GGUF, 3 shards / 1,628 tensors 83,274,984,576 bytes / 77.5559 GiB
Main tensor payload 83,263,928,920 bytes / 77.5456 GiB
FP8 PLE, 4 files / 128 logical parts 51,200,393,216 bytes / 47.6841 GiB
FP8 scale 2 bytes, BF16 0x3951
Total weight files, including scale 134,475,377,794 bytes / 125.2400 GiB

GGUF metadata/alignment, PLE cache, KV state and runtime workspaces are separate costs.

Quantization targets

Region Type
Routed gate/up, interior layers 2–45 Q4_K
Routed gate/up, edge layers 0, 1, 46, 47 Q5_K
Routed down, main 512 columns Q5_K
Routed down, 128-column tail Q5_0
MTP and most always-active matrices Primarily Q8_0
Non-quantizable matrices, norms and controls BF16 / F32 / I64
PLE n-gram table on SSD FP8 E4M3FN

The exact tensor map is in main-gguf-source-map.json. The main GGUF retains embedded MTP, vision tensors, integer PLE controls and the source chat template. The standalone tokenizer and processor assets are also included.

Published artifact layout

MQ-Q5-SSD-PLE-BF16/       Three main GGUF shards, recipe and checksums
PLE-FP8/                 Four FP8 embedding files, scale and manifest
reproduction/            Converter, tensor maps and verification reports

PLE provenance

All 128 Swift BF16 PLE parts match the Qwen BF16 reference. Complete pinned source-shard SHA-256 proves 124 parts; direct tensor-range SHA-256 proves the four parts in the two modified mixed shards. The identity report records both methods.

The official Qwen FP8 PLE was copied from the local Qwen model after that complete identity proof. It is the same FP8 sidecar used by the Baekpica Qwen3.8 and Uncensored releases. It is not a separate Swift FP8 checkpoint or a new BF16-to-FP8 quantization.

Download and use with ds4-dfm-rs

This is the dedicated qwen4exp SSD-PLE format. Use Baekpica/ds4-dfm-rs, with its FP8 PLE loader; generic GGUF runtime compatibility is not established.

swift_repo=Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
swift_root=./Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
hf download "$swift_repo" --local-dir "$swift_root"
(cd "$swift_root" && sha256sum -c SHA256SUMS)
export DS4_QWEN_PLE_DIR="$(realpath "$swift_root/PLE-FP8")"
./ds4 --cuda -m "$swift_root/MQ-Q5-SSD-PLE-BF16/Swift1.5-Qwen3.8-Flash-Next-MQ-Q5-SSD-PLE-BF16-00001-of-00003.gguf" \
  -c 8192 -n 128 -p "Explain binary search."

Set DS4_QWEN_PLE_DIR explicitly: this package contains FP8 PLE, while the main packaging name and input precision remain MQ-Q5-SSD-PLE-BF16.

Verification

Source-file checksums, all 1,658 source tensors, the 1,628-tensor Q5 recipe, PLE controls/layout, complete BF16 PLE identity, main GGUF structure and independent main/FP8 hashes passed. See artifact-manifest.json, SHA256SUMS and reproduction/.

After conversion, bounded DGX Spark CUDA serving checks passed with 262,144 configured context, two continuous banks, prefill 8,192, MTP draft 2, a 2 GiB FP8 PLE cache and 32 GiB SSD KV capacity. Text, simultaneous two-bank requests, prefix/tool continuation reuse and image input passed with zero memory faults. A full-length 262K prompt and cross-process disk restore were not tested. Swift throughput and task-quality evaluations have not been run; the graph above remains a Qwen reference, not a Swift measurement.

License and attribution

Swift Contribution: Swift Open License v1.0, Copyright 2026 UkisAI. Base model: Qwen Community License 1.0, Copyright 2026 Qwen. Both licenses and the original notices are included. NOTICE records the conversion changes. Quantizer code is MIT licensed; see ds4/LICENSE.

Thanks to UkisAI for Swift1.5, Qwen for the base model and FP8 PLE, GGML/llama.cpp for the quantization foundations, and the DwarfStar/ds4 contributors. This conversion is not an official UkisAI or Qwen release.