File size: 4,320 Bytes
09e211f
 
 
 
 
 
 
 
4a4e75f
09e211f
 
4a4e75f
09e211f
 
 
4a4e75f
 
09e211f
4a4e75f
 
 
09e211f
4a4e75f
09e211f
4a4e75f
09e211f
4a4e75f
 
 
09e211f
4a4e75f
09e211f
01d0ef7
09e211f
88b4502
4a4e75f
88b4502
 
 
 
 
09e211f
a48c4bf
 
09e211f
4a4e75f
09e211f
 
4a4e75f
09e211f
4a4e75f
09e211f
4a4e75f
 
 
 
88b4502
09e211f
 
4a4e75f
09e211f
4a4e75f
 
 
 
 
 
09e211f
4a4e75f
 
 
 
 
09e211f
4a4e75f
 
 
 
 
 
 
 
09e211f
4a4e75f
 
 
 
 
09e211f
4a4e75f
09e211f
 
4a4e75f
09e211f
4a4e75f
 
 
 
 
 
 
 
 
09e211f
c21ae99
 
 
 
4a4e75f
09e211f
4a4e75f
 
09e211f
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
---
tags:
- techwithsergiu
library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
pipeline_tag: text-generation
base_model:
- techwithsergiu/Qwen3.5-text-9B
---

# Qwen3.5-text-9B-bnb-4bit

<img width="400px" src="https://qianwen-res.oss-accelerate.aliyuncs.com/logo_qwen3.5.png">

BNB NF4 4-bit quantization of [techwithsergiu/Qwen3.5-text-9B](https://huggingface.co/techwithsergiu/Qwen3.5-text-9B) —
a text-only derivative of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B).

**No visual tower** — text input only. This is the recommended base for Unsloth LoRA
text fine-tuning: smaller VRAM footprint, no visual-dependency complexity, cleaner
adapter targeting.

Inference has been verified. LoRA fine-tuning docs are pending — see Fine-tuning section below.

## What was changed from the original Qwen3.5-9B

- Visual tower removed (same as `Qwen3.5-text-9B`)
- Text backbone quantized to BNB NF4 double-quant (`bnb_4bit_quant_type=nf4`, `bnb_4bit_compute_dtype=bfloat16`)
- `lm_head.weight` kept at **bf16** for output quality / stability

## Model family

![](diagrams/diagram_01.png)

| Model | Type | Base model |
|---|---|---|
| [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | f16 · VLM · source | — |
| [techwithsergiu/Qwen3.5-9B-bnb-4bit](https://huggingface.co/techwithsergiu/Qwen3.5-9B-bnb-4bit) | BNB NF4 · VLM | Qwen/Qwen3.5-9B |
| [techwithsergiu/Qwen3.5-text-9B](https://huggingface.co/techwithsergiu/Qwen3.5-text-9B) | bf16 · text-only | Qwen/Qwen3.5-9B |
| **[techwithsergiu/Qwen3.5-text-9B-bnb-4bit](https://huggingface.co/techwithsergiu/Qwen3.5-text-9B-bnb-4bit)** | BNB NF4 · text-only | Qwen3.5-text-9B |
| [techwithsergiu/Qwen3.5-text-9B-GGUF](https://huggingface.co/techwithsergiu/Qwen3.5-text-9B-GGUF) | GGUF quants | Qwen3.5-text-9B |

The visual tower scales with model size (~0.19 GB for 0.8B, ~0.62 GB for 2B/4B, ~0.85 GB for 9B).
BNB text-only models are roughly 34% of the original f16 size (4B example: 9.32 GB → 3.12 GB).

## Inference

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "techwithsergiu/Qwen3.5-text-9B-bnb-4bit"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    device_map="auto",
    trust_remote_code=True,
)

messages = [{"role": "user", "content": "What is the capital of Romania?"}]

# Thinking OFF — direct answer
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
response = tokenizer.decode(
    outputs[0][inputs["input_ids"].shape[1]:],
    skip_special_tokens=True,
)
print(response)

# Thinking ON — chain-of-thought before the answer
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=1024)
response = tokenizer.decode(
    outputs[0][inputs["input_ids"].shape[1]:],
    skip_special_tokens=True,
)
print(response)
```

## Fine-tuning

> **TBD** — LoRA training with this model has not been documented yet.
> The model has been verified for inference (text generation, thinking ON/OFF).
> The expectation is that standard Unsloth LoRA training applies — this is a
> text-only BNB 4-bit model architecturally identical to models Unsloth supports —
> but this has not been tested yet and there is no official Qwen3.5 text-only
> training guide to reference.
>
> For VLM (image + text) fine-tuning of the full model, see:
> [unsloth.ai/docs/models/qwen3.5/fine-tune](https://unsloth.ai/docs/models/qwen3.5/fine-tune)

## Pipeline diagram

![](diagrams/diagram_02.png)

## Acknowledgements

Based on [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)
by the Qwen Team. If you use this model in research, please cite the original:

```bibtex
@misc{qwen3.5,
    title  = {{Qwen3.5}: Towards Native Multimodal Agents},
    author = {{Qwen Team}},
    month  = {February},
    year   = {2026},
    url    = {https://qwen.ai/blog?id=qwen3.5}
}
```