Image-Text-to-Text
GGUF
qwen4_exp
mixed-quant
qwen4exp
qwen3.8-flash-next
swift1.5
ds4
dgx-spark
ssd-offload
conversational
Instructions to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Use Docker
docker model run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- LM Studio
- Jan
- vLLM
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- Ollama
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Ollama:
ollama run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- Unsloth Desktop
- Pi
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Docker Model Runner:
docker model run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- Lemonade
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Run and chat with the model
lemonade run user.Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF-BF16
List all available models
lemonade list
- Hermes Agent
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 7,226 Bytes
8ed94fb ae48755 8ed94fb ae48755 8ed94fb b090c63 8ed94fb ae48755 8ed94fb ae48755 8ed94fb ae48755 8ed94fb ae48755 8ed94fb ae48755 8ed94fb ae48755 8ed94fb b090c63 8ed94fb ae48755 8ed94fb ae48755 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | ---
license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/ukisai/Swift1.5-Qwen3.8-Flash-Next/blob/0bd4fe22431372cdad1979267d3ab45aa7e6150a/LICENSE
base_model: ukisai/Swift1.5-Qwen3.8-Flash-Next
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- gguf
- mixed-quant
- qwen4exp
- qwen3.8-flash-next
- swift1.5
- ds4
- dgx-spark
- ssd-offload
---
# Swift1.5-Qwen3.8-Flash-Next Mixed-Quant GGUF
> **One smaller Q5 compute backbone with an official FP8 SSD-PLE sidecar.**
> The main precision map matches the published Qwen3.8 and Uncensored
> `MQ-Q5-SSD-PLE-BF16` recipe. All 128 Swift BF16 PLE parts match Qwen,
> allowing the existing official FP8 PLE files to be reused unchanged.
## Performance reference: Qwen base Q5

*Reproduced unchanged from the [Qwen model card](https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-SSD-PLE-GGUF/blob/f92da0134498acfefe55aac7bd29e4dcdbbdbfe7/README.md).
Measured on Qwen base Q5, not Swift: one DGX Spark, 2 GiB PLE cache,
16 workers, MTP draft 2; 2,048-token incremental prefill and 128 greedy
tokens per frontier from 2K to 64K. Curves show per-frontier medians;
bands show the observed min–max across three fresh-process runs.*
Swift uses the same graph, tensor shapes, mixed-quant recipe and identical
FP8 PLE files, so similar runtime throughput is expected under matched
conditions. Equal speed has not been measured: routing, generated tokens,
PLE access patterns and MTP acceptance can differ with the main weights.
The BF16 curve is a Qwen comparison; this Swift package ships FP8 PLE only.
[Benchmark protocol, raw data and limits](https://github.com/Baekpica/ds4-dfm-rs/blob/407509a06c366e372499817397918d6690bedcc4/docs/qwen38-ple-fp8.md).
## Support my work
Support model conversion, inference optimization, profiling and open validation:
[Buy Me a Coffee](https://www.buymeacoffee.com/baekpica) · [GitHub Sponsors](https://github.com/sponsors/Baekpica).
Independent Baekpica conversion of
[ukisai/Swift1.5-Qwen3.8-Flash-Next](https://huggingface.co/ukisai/Swift1.5-Qwen3.8-Flash-Next/tree/0bd4fe22431372cdad1979267d3ab45aa7e6150a).
The main weights use the same smaller `MQ-Q5-SSD-PLE-BF16` tensor recipe as
the published Qwen3.8 and Uncensored SSD-PLE models. Conversion ran on CPU
from source BF16 weights, with no imatrix, pruning, expert dropping or layer dropping.
The memory hierarchy follows the existing
[Qwen3.8 mixed-quant SSD-PLE release](https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-SSD-PLE-GGUF):
the 51.2B-parameter PLE table resides on SSD with a bounded runtime page cache,
while the 128.8B-parameter compute backbone uses a mixed-precision GGUF.
## Variant and storage
| Artifact | Storage |
|---|---:|
| Main GGUF, 3 shards / 1,628 tensors | 83,274,984,576 bytes / 77.5559 GiB |
| Main tensor payload | 83,263,928,920 bytes / 77.5456 GiB |
| FP8 PLE, 4 files / 128 logical parts | 51,200,393,216 bytes / 47.6841 GiB |
| FP8 scale | 2 bytes, BF16 `0x3951` |
| Total weight files, including scale | 134,475,377,794 bytes / 125.2400 GiB |
GGUF metadata/alignment, PLE cache, KV state and runtime workspaces are separate costs.
## Quantization targets
| Region | Type |
|---|---|
| Routed gate/up, interior layers 2–45 | Q4_K |
| Routed gate/up, edge layers 0, 1, 46, 47 | Q5_K |
| Routed down, main 512 columns | Q5_K |
| Routed down, 128-column tail | Q5_0 |
| MTP and most always-active matrices | Primarily Q8_0 |
| Non-quantizable matrices, norms and controls | BF16 / F32 / I64 |
| PLE n-gram table on SSD | FP8 E4M3FN |
The exact tensor map is in
[`main-gguf-source-map.json`](reproduction/manifests/swift/ssd-ple-q5/main-gguf-source-map.json).
The main GGUF retains embedded MTP, vision tensors, integer PLE controls and
the source chat template. The standalone tokenizer and processor assets are also included.
## Published artifact layout
```text
MQ-Q5-SSD-PLE-BF16/ Three main GGUF shards, recipe and checksums
PLE-FP8/ Four FP8 embedding files, scale and manifest
reproduction/ Converter, tensor maps and verification reports
```
## PLE provenance
All 128 Swift BF16 PLE parts match the Qwen BF16 reference. Complete pinned
source-shard SHA-256 proves 124 parts; direct tensor-range SHA-256 proves the
four parts in the two modified mixed shards. The
[`identity report`](reproduction/manifests/swift/verify-ple-identity.json) records both methods.
The official [Qwen FP8 PLE](https://huggingface.co/Qwen/Qwen3.8-Flash-Next-FP8/tree/236dfdf285828023ca3bcd3f37366c58a3469b13)
was copied from the local Qwen model after that complete identity proof.
It is the same FP8 sidecar used by the Baekpica Qwen3.8 and Uncensored releases.
It is not a separate Swift FP8 checkpoint or a new BF16-to-FP8 quantization.
## Download and use with ds4-dfm-rs
This is the dedicated `qwen4exp` SSD-PLE format. Use
[`Baekpica/ds4-dfm-rs`](https://github.com/Baekpica/ds4-dfm-rs), with its FP8 PLE loader;
generic GGUF runtime compatibility is not established.
```bash
swift_repo=Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
swift_root=./Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
hf download "$swift_repo" --local-dir "$swift_root"
(cd "$swift_root" && sha256sum -c SHA256SUMS)
export DS4_QWEN_PLE_DIR="$(realpath "$swift_root/PLE-FP8")"
./ds4 --cuda -m "$swift_root/MQ-Q5-SSD-PLE-BF16/Swift1.5-Qwen3.8-Flash-Next-MQ-Q5-SSD-PLE-BF16-00001-of-00003.gguf" \
-c 8192 -n 128 -p "Explain binary search."
```
Set `DS4_QWEN_PLE_DIR` explicitly: this package contains FP8 PLE, while the
main packaging name and input precision remain `MQ-Q5-SSD-PLE-BF16`.
## Verification
Source-file checksums, all 1,658 source tensors, the 1,628-tensor Q5 recipe,
PLE controls/layout, complete BF16 PLE identity, main GGUF structure and
independent main/FP8 hashes passed. See [`artifact-manifest.json`](artifact-manifest.json),
[`SHA256SUMS`](SHA256SUMS) and [`reproduction/`](reproduction/README.md).
After conversion, bounded DGX Spark CUDA serving checks passed with 262,144
configured context, two continuous banks, prefill 8,192, MTP draft 2, a 2 GiB
FP8 PLE cache and 32 GiB SSD KV capacity. Text, simultaneous two-bank requests,
prefix/tool continuation reuse and image input passed with zero memory faults.
A full-length 262K prompt and cross-process disk restore were not tested.
Swift throughput and task-quality evaluations have not been run; the graph
above remains a Qwen reference, not a Swift measurement.
## License and attribution
Swift Contribution: [Swift Open License v1.0](LICENSE), Copyright 2026 UkisAI.
Base model: [Qwen Community License 1.0](LICENSE-QWEN), Copyright 2026 Qwen.
Both licenses and the original notices are included. [`NOTICE`](NOTICE) records
the conversion changes. Quantizer code is MIT licensed; see [`ds4/LICENSE`](ds4/LICENSE).
Thanks to UkisAI for Swift1.5, Qwen for the base model and FP8 PLE, GGML/llama.cpp
for the quantization foundations, and the DwarfStar/ds4 contributors.
This conversion is not an official UkisAI or Qwen release.
|