techwithsergiu's picture
docs: upd diagrams, new qwen-qlora-train, qwen35-toolkit links
d10f238 verified
|
Raw
History Blame
5.02 kB
---
tags:
- techwithsergiu
- gguf
- qwen3_5_text
library_name: gguf
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-0.8B/blob/main/LICENSE
pipeline_tag: text-generation
base_model:
- techwithsergiu/Qwen3.5-text-0.8B
---
# Qwen3.5-text-0.8B-GGUF
<img width="400px" src="https://qianwen-res.oss-accelerate.aliyuncs.com/logo_qwen3.5.png">
GGUF quants of [techwithsergiu/Qwen3.5-text-0.8B](https://huggingface.co/techwithsergiu/Qwen3.5-text-0.8B) β€”
the text-only bf16 derivative of [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B).
The visual tower has been removed before conversion. All text-backbone weights are
**identical** to the original β€” no retraining, no weight changes, no quality loss for
text tasks.
## Quants
| File | Type | Size | Notes |
|---|---|---|---|
| `Qwen3.5-text-0.8B-Q8_0.gguf` | Q8_0 | ~53% of f16 | near-lossless β€” for high-quality inference |
| `Qwen3.5-text-0.8B-Q6_K.gguf` | Q6_K | ~41% of f16 | excellent quality, good balance with f16 |
| `Qwen3.5-text-0.8B-Q5_K_M.gguf` | Q5_K_M | ~37% of f16 | very good quality, smaller than Q6 |
| `Qwen3.5-text-0.8B-Q4_K_M.gguf` | Q4_K_M | ~31% of f16 | βœ… recommended β€” best size/quality balance |
| `Qwen3.5-text-0.8B-Q4_K_S.gguf` | Q4_K_S | ~30% of f16 | optional β€” slightly smaller, slightly lower quality |
## Model family
![](diagrams/diagram_01.png)
| Model | Type | Base model |
|---|---|---|
| [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) | f16 Β· VLM Β· source | β€” |
| [techwithsergiu/Qwen3.5-0.8B-bnb-4bit](https://huggingface.co/techwithsergiu/Qwen3.5-0.8B-bnb-4bit) | BNB NF4 Β· VLM | Qwen/Qwen3.5-0.8B |
| [techwithsergiu/Qwen3.5-text-0.8B](https://huggingface.co/techwithsergiu/Qwen3.5-text-0.8B) | bf16 Β· text-only | Qwen/Qwen3.5-0.8B |
| [techwithsergiu/Qwen3.5-text-0.8B-bnb-4bit](https://huggingface.co/techwithsergiu/Qwen3.5-text-0.8B-bnb-4bit) | BNB NF4 Β· text-only | Qwen3.5-text-0.8B |
| **[techwithsergiu/Qwen3.5-text-0.8B-GGUF](https://huggingface.co/techwithsergiu/Qwen3.5-text-0.8B-GGUF)** | GGUF quants | Qwen3.5-text-0.8B |
The GGUF repo is derived from the text-only f16 model β€” same weights, different container
format. `base_model` points to the f16 text variant to keep the VLM and text lineages
distinct on the Hub.
## Inference
### llama.cpp
```bash
./llama.cpp/build/bin/llama-cli \
-m Qwen3.5-text-0.8B-Q4_K_M.gguf \
-p "What is the capital of Romania?" \
-n 256
```
### LM Studio
Load any `.gguf` file from this repo directly in [LM Studio](https://lmstudio.ai).
Recommended quant: `Q4_K_M`.
### Thinking mode
Qwen3.5 supports an optional chain-of-thought `<think>` block before the answer.
Thinking is **enabled by default** in llama.cpp.
**Note:** `--chat-template-kwargs '{"enable_thinking":...}'` is deprecated β€” do not use.
**Known issue:** `--reasoning off` is accepted but does not actually disable thinking.
**Workaround:** use `--reasoning-budget 0` β€” this reliably disables the `<think>` block.
Track the bug at [llama.cpp issues](https://github.com/ggml-org/llama.cpp/issues).
```bash
# Thinking OFF β€” direct answer (workaround: --reasoning-budget 0)
./llama.cpp/build/bin/llama-cli \
-m Qwen3.5-text-0.8B-Q4_K_M.gguf \
--reasoning-budget 0 \
-p "What is the capital of Romania?" \
-n 256
# Thinking ON β€” default, no flag needed
./llama.cpp/build/bin/llama-cli \
-m Qwen3.5-text-0.8B-Q4_K_M.gguf \
-p "What is 17 Γ— 34?" \
-n 1024
```
## Pipeline diagram
![](diagrams/diagram_02.png)
## From fine-tuned adapter to GGUF
If you have a LoRA adapter trained with
[qwen-qlora-train](https://techwithsergiu.github.io/qwen-qlora-train),
merge it first, then convert to GGUF:
```bash
# 1. Merge adapter into f16 weights
qlora-merge \
--base Qwen/Qwen3.5-0.8B \
--adapter adapters/<run_name> \
--output merged/qwen35-text-0.8B-sft-f16
# 2. Convert merged model to GGUF (requires llama.cpp)
python llama.cpp/convert_hf_to_gguf.py merged/qwen35-text-0.8B-sft-f16 \
--outtype f16 \
--outfile merged/qwen35-text-0.8B-sft-F16.gguf
# 3. Quantize
./llama.cpp/build/bin/llama-quantize \
merged/qwen35-text-0.8B-sft-F16.gguf \
merged/qwen35-text-0.8B-sft-Q4_K_M.gguf \
Q4_K_M
```
Full post-training workflow is documented in
[qwen-qlora-train β†’ Post-merge workflow](https://techwithsergiu.github.io/qwen-qlora-train/post-merge-workflow.html).
## Conversion
Converted using [qwen35-toolkit](https://techwithsergiu.github.io/qwen35-toolkit) β€”
a Python toolkit for BNB quantization, visual tower removal, verification and
HF Hub publishing of Qwen3.5 models.
---
## Acknowledgements
Based on [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B)
by the Qwen Team. If you use this model in research, please cite the original:
```bibtex
@misc{qwen3.5,
title = {{Qwen3.5}: Towards Native Multimodal Agents},
author = {{Qwen Team}},
month = {February},
year = {2026},
url = {https://qwen.ai/blog?id=qwen3.5}
}
```