Instructions to use nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi") model = AutoModelForMultimodalLM.from_pretrained("nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi
- SGLang
How to use nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi with Docker Model Runner:
docker model run hf.co/nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi
- Qwen3-VL-4B-Thinking-Pashto-Zi 🇦🇫
- 🌍 Why Token Surgery on Qwen3-VL?
- 🔬 Tokenizer Audit
- ➕ New Vocabulary
- 🧬 Embedding Initialization
- 🔎 Embedding Forensics
- 📊 Embedding Distribution
- 🧠 Consistency Across Qwen3-VL Model Sizes
- 🔐 Forensic Gate
- 🧪 Verification
- 🧠 Important Note
- 🛠️ Intended Use
- 📦 Repository Contents
- 📋 Base Model
- 🚀 Loading the Model
- 🔤 Testing Newly Added Characters
- 📈 Token Efficiency
- ⚠️ Status
- 📜 Forensic Report
- 👤 Author
- 📄 License
- ⭐ Project
Qwen3-VL-4B-Thinking-Pashto-Zi 🇦🇫
Pashto + Urdu + Persian + Balochi + Sindhi Token Surgery for Qwen/Qwen3-VL-4B-Thinking
This repository contains a tokenizer and embedding-extension version of:
Qwen/Qwen3-VL-4B-Thinking
with additional single-character tokens for Pashto, Sindhi, Balochi, and related Arabic-script characters that were still missing from the tokenizer — plus the full Eastern Arabic-Indic digit set and several bidi / harakat controls.
Model
- Base model:
Qwen/Qwen3-VL-4B-Thinking - Model repository:
nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi - Original tokenizer vocabulary:
151,669 - Extended vocabulary:
151,700 - New tokens:
31 - Embedding dimension:
2560 - Embedding dtype during surgery:
torch.bfloat16
The original embedding matrix had the shape:
(151936, 2560)
After tokenizer extension and resizing, the language-model embedding matrix has the shape:
(151700, 2560)
The input and output embeddings share the same storage (tied embeddings).
The vision tower of Qwen3-VL is not vocab-indexed and was not touched.
🌍 Why Token Surgery on Qwen3-VL?
Qwen3-VL already ships with very broad Arabic-script coverage. Of the 175 unique atoms audited, 144 were already single tokens — including all Pashto-specific letters, most Sindhi letters, and most Balochi vowels.
But 31 atoms were still missing and split into byte-fallback pairs. For Pashto, Sindhi, Balochi, and Persian/Urdu numeric text, this is a real tokenization inefficiency.
Examples of atoms that were previously split:
ٽ
ڇ
ڏ
ڙ
ڦ
ڻ
ۏ
۰
۱
۲
۳
۴
۵
۶
۷
۸
۹
These characters previously tokenized as multi-byte fragments such as
[151, 108], [150, 237], etc.
The purpose of this project is to give them their own independent vocabulary entries.
🔬 Tokenizer Audit
A total of 175 unique atoms were audited (181 raw entries, with intentional duplicates preserved for a separate "push-all" script).
Existing single-token atoms : 144
Atoms requiring new token : 31
The audit found 31 missing atoms requiring new vocabulary entries.
Examples of previously missing characters:
ٽ
ڇ
ڏ
ڙ
ڦ
ڻ
ۏ
Additional extended Arabic-Indic digits, bidi controls, and harakat were also added.
➕ New Vocabulary
The original vocabulary contained:
151669
After adding 31 new tokens:
151700
Exact New Token ID Map
| Token | ID |
|---|---|
| ٽ | 151669 |
| ڇ | 151670 |
| ڏ | 151671 |
| ڙ | 151672 |
| ڦ | 151673 |
| ڻ | 151674 |
| ۏ | 151675 |
| ۰ | 151676 |
| ۱ | 151677 |
| ۲ | 151678 |
| ۳ | 151679 |
| ۴ | 151680 |
| ۵ | 151681 |
| ۶ | 151682 |
| ۷ | 151683 |
| ۸ | 151684 |
| ۹ | 151685 |
| ؉ | 151686 |
| U+200D | 151687 |
| U+200F | 151688 |
| U+202A | 151689 |
| U+202B | 151690 |
| U+202C | 151691 |
| U+202D | 151692 |
| U+202E | 151693 |
| ٌ | 151694 |
| ٍ | 151695 |
| ٓ | 151696 |
| ٔ | 151697 |
| ٕ | 151698 |
| ٰ | 151699 |
🧬 Embedding Initialization
Each of the 31 new vocabulary entries receives its own independent embedding row.
The new embeddings were initialized independently with:
Initializer range: 0.02
Embedding dimension: 2560
The forensic audit confirmed:
Expected shape: (31, 2560)
Actual shape: (31, 2560)
Each new token therefore owns an independent 2560-dimensional embedding vector.
🔎 Embedding Forensics
An exact duplicate search was performed across the new embedding rows.
Result:
NO EXACT DUPLICATE EMBEDDING ROWS
A pairwise cosine similarity analysis was also performed.
The maximum cosine similarity between distinct new-token embeddings was:
0.046955175698
The corresponding pair was:
۲ <-> U+202E
This confirms that the newly initialized embedding rows are distinct rather than duplicated copies.
📊 Embedding Distribution
Original vocabulary
mean = -0.0000246564
std = 0.0215103794
norm mean = 1.07593632
New tokens
mean = -0.0000160644
std = 0.0199598726
norm mean = 1.00979233
The new-token embeddings were independently initialized and have their own embedding distribution. Their average norm sits very close to the pretrained rows, so continued pretraining has only a small distribution gap to close.
🧠 Consistency Across Qwen3-VL Model Sizes
The same surgery on Qwen/Qwen3-VL-2B-Thinking also produced exactly 31 new
tokens, with the same IDs. This confirms that Qwen3-VL uses a single shared
tokenizer across the 2B and 4B sizes, and that the "Arabic-script gap" is
identical regardless of parameter count.
| Base model | Embed dim | New tokens | Final vocab |
|---|---|---|---|
| Qwen3-VL-2B-Thinking | 2048 | 31 | 151,700 |
| Qwen3-VL-4B-Thinking | 2560 | 31 | 151,700 |
🔐 Forensic Gate
The final forensic checks confirmed:
✓ Input/output embeddings share storage (tied)
✓ Every new atom owns an independent embedding row
✓ No exact duplicate embedding rows
✓ No tokenizer-ID/indexing corruption detected
✓ Embedding matrix successfully resized
✓ Vision tower untouched (not vocab-indexed)
✓ Model successfully saved
✓ Tokenizer successfully saved
✓ Model successfully loaded again from Hugging Face
🧪 Verification
After uploading the model, it was loaded again from:
nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi
The loaded model reported:
Vocab size:
151700
Embedding shape:
torch.Size([151700, 2560])
The newly added tokens were also verified after reloading.
Examples:
ٽ -> 151669
ڇ -> 151670
ڏ -> 151671
ڙ -> 151672
ڦ -> 151673
ڻ -> 151674
Existing tokens that were already present in the original tokenizer retain their original IDs.
For example:
ښ -> 147996
ځ -> 148519
ږ -> 150389
ټ -> 148249
ډ -> 148967
ړ -> 148871
ګ -> 150392
These are existing tokenizer entries, not newly added vocabulary entries.
🧠 Important Note
This repository represents a tokenizer and embedding vocabulary extension.
The 31 newly added token embeddings were freshly initialized during the vocabulary surgery.
This does not by itself mean that the model has already learned full Pashto (or Sindhi, Balochi, …) language knowledge for these new atoms.
For the newly introduced tokens to acquire meaningful linguistic representations, further training such as continued pretraining / causal language modeling on multilingual Arabic-script data is required.
In other words:
Tokenizer surgery
↓
31 new Pashto + Sindhi + Balochi + Arabic-script atoms
↓
Independent embedding rows
↓
Continued pretraining
↓
Learned representations
🛠️ Intended Use
This model is intended as an experimental foundation for further Arabic-script language-model development on top of Qwen3-VL.
Potential next steps include:
- Pashto continued pretraining
- Urdu continued pretraining
- Sindhi continued pretraining
- Balochi continued pretraining
- Multilingual Arabic-script causal language modeling
- Vision-language instruction tuning
- OCR experiments on Arabic-script documents
- Tokenizer evaluation
- Token efficiency evaluation
- Language generation experiments
📦 Repository Contents
The repository contains the model, tokenizer, configuration, generation configuration, chat template, and forensic report.
Important files include:
config.json
model.safetensors
generation_config.json
tokenizer_config.json
tokenizer.json
vocab.json
merges.txt
chat_template.jinja
pashto_urdu_fresh_embedding_forensics_Qwen_Qwen3-VL-4B-Thinking.json
The forensic report records the tokenizer audit, vocabulary additions, embedding checks, and verification results.
📋 Base Model
This project is based on:
Qwen/Qwen3-VL-4B-Thinking
Base model:
Qwen — Qwen3-VL-4B-Thinking
This repository does not claim to reproduce or replace the original base model. It provides an extended tokenizer vocabulary and corresponding resized embedding matrix.
🚀 Loading the Model
from transformers import AutoTokenizer, Qwen3VLForConditionalGeneration
model_id = "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = Qwen3VLForConditionalGeneration.from_pretrained(
model_id,
dtype="auto",
device_map="auto"
)
print("Vocabulary:", len(tokenizer))
print("Embedding:", model.get_input_embeddings().weight.shape)
Expected output:
Vocabulary: 151700
Embedding: torch.Size([151700, 2560])
🔤 Testing Newly Added Characters
You can inspect the newly added atoms with:
new_atoms = [
# Sindhi
"ٽ", "ڇ", "ڏ", "ڙ", "ڦ", "ڻ",
# Balochi
"ۏ",
# Extended Arabic-Indic digits
"۰", "۱", "۲", "۳", "۴", "۵", "۶", "۷", "۸", "۹",
# Symbol
"؉",
# Bidi controls
"\u200d", "\u200f",
"\u202a", "\u202b", "\u202c", "\u202d", "\u202e",
# Harakat
"ٌ", "ٍ", "ٓ", "ٔ", "ٕ", "ٰ",
]
for char in new_atoms:
token_id = tokenizer.convert_tokens_to_ids(char)
print(repr(char), "->", token_id)
Expected IDs run from 151669 through 151699 (see the table above).
📈 Token Efficiency
After the surgery, the following text categories tokenize more efficiently:
| Category | Before | After |
|---|---|---|
Eastern Arabic-Indic digits (۰۱۲۳۴۵۶۷۸۹) |
2 tokens each | 1 token each |
Sindhi consonants ٽ ڇ ڏ ڙ ڦ ڻ |
2 tokens each | 1 token each |
Balochi vowel ۏ |
2 tokens | 1 token |
| Bidi controls | 2 tokens each | 1 token each |
| Select harakat | 2 tokens each | 1 token each |
Approximate efficiency gains on Pashto, Sindhi, Balochi, and Persian/Urdu numeric text: ~10–20% fewer tokens for text that uses these atoms.
⚠️ Status
Tokenizer surgery: ✅ Complete
Vocabulary resize: ✅ Complete
Embedding forensic audit: ✅ Passed
Model save: ✅ Complete
Tokenizer save: ✅ Complete
Hugging Face upload: ✅ Complete
Reload verification: ✅ Complete
Continued pretraining: ⏳ Not performed as part of this surgery
SFT / instruction tuning: ⏳ Not performed as part of this surgery
📜 Forensic Report
A complete forensic report was generated during the surgery:
pashto_urdu_fresh_embedding_forensics_Qwen_Qwen3-VL-4B-Thinking.json
The report contains:
- tokenizer audit
- missing-token inventory
- new token IDs
- embedding shapes
- embedding fingerprints
- duplicate detection
- pairwise cosine analysis
- old/new embedding distribution
- final forensic gate
- model verification
👤 Author
nassimjp
Pashto AI / Pashto NLP experiments
📄 License
Please refer to the license of the original base model and repository configuration before redistributing modified model weights.
⭐ Project
Qwen3-VL-4B-Thinking-Pashto-Zi
A focused tokenizer surgery on top of an already multilingual base:
Qwen/Qwen3-VL-4B-Thinking
+
31 new Pashto + Sindhi + Balochi + Arabic-script atoms
↓
151,669 → 151,700 vocabulary
↓
Multilingual Arabic-script tokenizer foundation
+
Vision tower preserved intact
The tokenizer now has dedicated vocabulary entries for the missing Pashto, Sindhi, Balochi, and related Arabic-script characters that Qwen3-VL previously split into byte fragments.
---
- Downloads last month
- 35
Model tree for nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi
Base model
Qwen/Qwen3-VL-4B-Thinking