Instructions to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Use Docker
docker model run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- LM Studio
- Jan
- vLLM
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- Ollama
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Ollama:
ollama run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- Unsloth Desktop
- Pi
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Docker Model Runner:
docker model run hf.co/Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
- Lemonade
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Run and chat with the model
lemonade run user.Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF-BF16
List all available models
lemonade list
- Hermes Agent
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Swift1.5-Qwen3.8-Flash-Next Mixed-Quant GGUF
One smaller Q5 compute backbone with an official FP8 SSD-PLE sidecar. The main precision map matches the published Qwen3.8 and Uncensored
MQ-Q5-SSD-PLE-BF16recipe. All 128 Swift BF16 PLE parts match Qwen, allowing the existing official FP8 PLE files to be reused unchanged.
Performance reference: Qwen base Q5
Reproduced unchanged from the Qwen model card. Measured on Qwen base Q5, not Swift: one DGX Spark, 2 GiB PLE cache, 16 workers, MTP draft 2; 2,048-token incremental prefill and 128 greedy tokens per frontier from 2K to 64K. Curves show per-frontier medians; bands show the observed min–max across three fresh-process runs.
Swift uses the same graph, tensor shapes, mixed-quant recipe and identical FP8 PLE files, so similar runtime throughput is expected under matched conditions. Equal speed has not been measured: routing, generated tokens, PLE access patterns and MTP acceptance can differ with the main weights. The BF16 curve is a Qwen comparison; this Swift package ships FP8 PLE only. Benchmark protocol, raw data and limits.
Support my work
Support model conversion, inference optimization, profiling and open validation: Buy Me a Coffee · GitHub Sponsors.
Independent Baekpica conversion of
ukisai/Swift1.5-Qwen3.8-Flash-Next.
The main weights use the same smaller MQ-Q5-SSD-PLE-BF16 tensor recipe as
the published Qwen3.8 and Uncensored SSD-PLE models. Conversion ran on CPU
from source BF16 weights, with no imatrix, pruning, expert dropping or layer dropping.
The memory hierarchy follows the existing Qwen3.8 mixed-quant SSD-PLE release: the 51.2B-parameter PLE table resides on SSD with a bounded runtime page cache, while the 128.8B-parameter compute backbone uses a mixed-precision GGUF.
Variant and storage
| Artifact | Storage |
|---|---|
| Main GGUF, 3 shards / 1,628 tensors | 83,274,984,576 bytes / 77.5559 GiB |
| Main tensor payload | 83,263,928,920 bytes / 77.5456 GiB |
| FP8 PLE, 4 files / 128 logical parts | 51,200,393,216 bytes / 47.6841 GiB |
| FP8 scale | 2 bytes, BF16 0x3951 |
| Total weight files, including scale | 134,475,377,794 bytes / 125.2400 GiB |
GGUF metadata/alignment, PLE cache, KV state and runtime workspaces are separate costs.
Quantization targets
| Region | Type |
|---|---|
| Routed gate/up, interior layers 2–45 | Q4_K |
| Routed gate/up, edge layers 0, 1, 46, 47 | Q5_K |
| Routed down, main 512 columns | Q5_K |
| Routed down, 128-column tail | Q5_0 |
| MTP and most always-active matrices | Primarily Q8_0 |
| Non-quantizable matrices, norms and controls | BF16 / F32 / I64 |
| PLE n-gram table on SSD | FP8 E4M3FN |
The exact tensor map is in
main-gguf-source-map.json.
The main GGUF retains embedded MTP, vision tensors, integer PLE controls and
the source chat template. The standalone tokenizer and processor assets are also included.
Published artifact layout
MQ-Q5-SSD-PLE-BF16/ Three main GGUF shards, recipe and checksums
PLE-FP8/ Four FP8 embedding files, scale and manifest
reproduction/ Converter, tensor maps and verification reports
PLE provenance
All 128 Swift BF16 PLE parts match the Qwen BF16 reference. Complete pinned
source-shard SHA-256 proves 124 parts; direct tensor-range SHA-256 proves the
four parts in the two modified mixed shards. The
identity report records both methods.
The official Qwen FP8 PLE was copied from the local Qwen model after that complete identity proof. It is the same FP8 sidecar used by the Baekpica Qwen3.8 and Uncensored releases. It is not a separate Swift FP8 checkpoint or a new BF16-to-FP8 quantization.
Download and use with ds4-dfm-rs
This is the dedicated qwen4exp SSD-PLE format. Use
Baekpica/ds4-dfm-rs, with its FP8 PLE loader;
generic GGUF runtime compatibility is not established.
swift_repo=Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
swift_root=./Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
hf download "$swift_repo" --local-dir "$swift_root"
(cd "$swift_root" && sha256sum -c SHA256SUMS)
export DS4_QWEN_PLE_DIR="$(realpath "$swift_root/PLE-FP8")"
./ds4 --cuda -m "$swift_root/MQ-Q5-SSD-PLE-BF16/Swift1.5-Qwen3.8-Flash-Next-MQ-Q5-SSD-PLE-BF16-00001-of-00003.gguf" \
-c 8192 -n 128 -p "Explain binary search."
Set DS4_QWEN_PLE_DIR explicitly: this package contains FP8 PLE, while the
main packaging name and input precision remain MQ-Q5-SSD-PLE-BF16.
Verification
Source-file checksums, all 1,658 source tensors, the 1,628-tensor Q5 recipe,
PLE controls/layout, complete BF16 PLE identity, main GGUF structure and
independent main/FP8 hashes passed. See artifact-manifest.json,
SHA256SUMS and reproduction/.
After conversion, bounded DGX Spark CUDA serving checks passed with 262,144 configured context, two continuous banks, prefill 8,192, MTP draft 2, a 2 GiB FP8 PLE cache and 32 GiB SSD KV capacity. Text, simultaneous two-bank requests, prefix/tool continuation reuse and image input passed with zero memory faults. A full-length 262K prompt and cross-process disk restore were not tested. Swift throughput and task-quality evaluations have not been run; the graph above remains a Qwen reference, not a Swift measurement.
License and attribution
Swift Contribution: Swift Open License v1.0, Copyright 2026 UkisAI.
Base model: Qwen Community License 1.0, Copyright 2026 Qwen.
Both licenses and the original notices are included. NOTICE records
the conversion changes. Quantizer code is MIT licensed; see ds4/LICENSE.
Thanks to UkisAI for Swift1.5, Qwen for the base model and FP8 PLE, GGML/llama.cpp for the quantization foundations, and the DwarfStar/ds4 contributors. This conversion is not an official UkisAI or Qwen release.
- Downloads last month
- 1,235
16-bit
Model tree for Baekpica/Swift1.5-Qwen3.8-Flash-Next-Mixed-Quant-GGUF
Base model
Qwen/Qwen3.8-Flash-Next