Qwen3-VL-4B-Thinking-Pashto-Zi 🇦🇫

Pashto + Urdu + Persian + Balochi + Sindhi Token Surgery for Qwen/Qwen3-VL-4B-Thinking

This repository contains a tokenizer and embedding-extension version of:

Qwen/Qwen3-VL-4B-Thinking

with additional single-character tokens for Pashto, Sindhi, Balochi, and related Arabic-script characters that were still missing from the tokenizer — plus the full Eastern Arabic-Indic digit set and several bidi / harakat controls.

Model

  • Base model: Qwen/Qwen3-VL-4B-Thinking
  • Model repository: nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi
  • Original tokenizer vocabulary: 151,669
  • Extended vocabulary: 151,700
  • New tokens: 31
  • Embedding dimension: 2560
  • Embedding dtype during surgery: torch.bfloat16

The original embedding matrix had the shape:

(151936, 2560)

After tokenizer extension and resizing, the language-model embedding matrix has the shape:

(151700, 2560)

The input and output embeddings share the same storage (tied embeddings).

The vision tower of Qwen3-VL is not vocab-indexed and was not touched.


🌍 Why Token Surgery on Qwen3-VL?

Qwen3-VL already ships with very broad Arabic-script coverage. Of the 175 unique atoms audited, 144 were already single tokens — including all Pashto-specific letters, most Sindhi letters, and most Balochi vowels.

But 31 atoms were still missing and split into byte-fallback pairs. For Pashto, Sindhi, Balochi, and Persian/Urdu numeric text, this is a real tokenization inefficiency.

Examples of atoms that were previously split:

ٽ
ڇ
ڏ
ڙ
ڦ
ڻ
ۏ
۰
۱
۲
۳
۴
۵
۶
۷
۸
۹

These characters previously tokenized as multi-byte fragments such as [151, 108], [150, 237], etc.

The purpose of this project is to give them their own independent vocabulary entries.


🔬 Tokenizer Audit

A total of 175 unique atoms were audited (181 raw entries, with intentional duplicates preserved for a separate "push-all" script).

Existing single-token atoms : 144
Atoms requiring new token   :  31

The audit found 31 missing atoms requiring new vocabulary entries.

Examples of previously missing characters:

ٽ
ڇ
ڏ
ڙ
ڦ
ڻ
ۏ

Additional extended Arabic-Indic digits, bidi controls, and harakat were also added.


➕ New Vocabulary

The original vocabulary contained:

151669

After adding 31 new tokens:

151700

Exact New Token ID Map

Token ID
ٽ 151669
ڇ 151670
ڏ 151671
ڙ 151672
ڦ 151673
ڻ 151674
ۏ 151675
۰ 151676
۱ 151677
۲ 151678
۳ 151679
۴ 151680
۵ 151681
۶ 151682
۷ 151683
۸ 151684
۹ 151685
؉ 151686
U+200D 151687
U+200F 151688
U+202A 151689
U+202B 151690
U+202C 151691
U+202D 151692
U+202E 151693
ٌ 151694
ٍ 151695
ٓ 151696
ٔ 151697
ٕ 151698
ٰ 151699

🧬 Embedding Initialization

Each of the 31 new vocabulary entries receives its own independent embedding row.

The new embeddings were initialized independently with:

Initializer range: 0.02
Embedding dimension: 2560

The forensic audit confirmed:

Expected shape: (31, 2560)
Actual shape:   (31, 2560)

Each new token therefore owns an independent 2560-dimensional embedding vector.


🔎 Embedding Forensics

An exact duplicate search was performed across the new embedding rows.

Result:

NO EXACT DUPLICATE EMBEDDING ROWS

A pairwise cosine similarity analysis was also performed.

The maximum cosine similarity between distinct new-token embeddings was:

0.046955175698

The corresponding pair was:

۲  <->  U+202E

This confirms that the newly initialized embedding rows are distinct rather than duplicated copies.


📊 Embedding Distribution

Original vocabulary

mean      = -0.0000246564
std       =  0.0215103794
norm mean =  1.07593632

New tokens

mean      = -0.0000160644
std       =  0.0199598726
norm mean =  1.00979233

The new-token embeddings were independently initialized and have their own embedding distribution. Their average norm sits very close to the pretrained rows, so continued pretraining has only a small distribution gap to close.


🧠 Consistency Across Qwen3-VL Model Sizes

The same surgery on Qwen/Qwen3-VL-2B-Thinking also produced exactly 31 new tokens, with the same IDs. This confirms that Qwen3-VL uses a single shared tokenizer across the 2B and 4B sizes, and that the "Arabic-script gap" is identical regardless of parameter count.

Base model Embed dim New tokens Final vocab
Qwen3-VL-2B-Thinking 2048 31 151,700
Qwen3-VL-4B-Thinking 2560 31 151,700

🔐 Forensic Gate

The final forensic checks confirmed:

✓ Input/output embeddings share storage (tied)
✓ Every new atom owns an independent embedding row
✓ No exact duplicate embedding rows
✓ No tokenizer-ID/indexing corruption detected
✓ Embedding matrix successfully resized
✓ Vision tower untouched (not vocab-indexed)
✓ Model successfully saved
✓ Tokenizer successfully saved
✓ Model successfully loaded again from Hugging Face

🧪 Verification

After uploading the model, it was loaded again from:

nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi

The loaded model reported:

Vocab size:
151700

Embedding shape:
torch.Size([151700, 2560])

The newly added tokens were also verified after reloading.

Examples:

ٽ -> 151669
ڇ -> 151670
ڏ -> 151671
ڙ -> 151672
ڦ -> 151673
ڻ -> 151674

Existing tokens that were already present in the original tokenizer retain their original IDs.

For example:

ښ -> 147996
ځ -> 148519
ږ -> 150389
ټ -> 148249
ډ -> 148967
ړ -> 148871
ګ -> 150392

These are existing tokenizer entries, not newly added vocabulary entries.


🧠 Important Note

This repository represents a tokenizer and embedding vocabulary extension.

The 31 newly added token embeddings were freshly initialized during the vocabulary surgery.

This does not by itself mean that the model has already learned full Pashto (or Sindhi, Balochi, …) language knowledge for these new atoms.

For the newly introduced tokens to acquire meaningful linguistic representations, further training such as continued pretraining / causal language modeling on multilingual Arabic-script data is required.

In other words:

Tokenizer surgery
        ↓
31 new Pashto + Sindhi + Balochi + Arabic-script atoms
        ↓
Independent embedding rows
        ↓
Continued pretraining
        ↓
Learned representations

🛠️ Intended Use

This model is intended as an experimental foundation for further Arabic-script language-model development on top of Qwen3-VL.

Potential next steps include:

  • Pashto continued pretraining
  • Urdu continued pretraining
  • Sindhi continued pretraining
  • Balochi continued pretraining
  • Multilingual Arabic-script causal language modeling
  • Vision-language instruction tuning
  • OCR experiments on Arabic-script documents
  • Tokenizer evaluation
  • Token efficiency evaluation
  • Language generation experiments

📦 Repository Contents

The repository contains the model, tokenizer, configuration, generation configuration, chat template, and forensic report.

Important files include:

config.json
model.safetensors
generation_config.json
tokenizer_config.json
tokenizer.json
vocab.json
merges.txt
chat_template.jinja

pashto_urdu_fresh_embedding_forensics_Qwen_Qwen3-VL-4B-Thinking.json

The forensic report records the tokenizer audit, vocabulary additions, embedding checks, and verification results.


📋 Base Model

This project is based on:

Qwen/Qwen3-VL-4B-Thinking

Base model:

Qwen — Qwen3-VL-4B-Thinking

This repository does not claim to reproduce or replace the original base model. It provides an extended tokenizer vocabulary and corresponding resized embedding matrix.


🚀 Loading the Model

from transformers import AutoTokenizer, Qwen3VLForConditionalGeneration

model_id = "nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi"

tokenizer = AutoTokenizer.from_pretrained(model_id)

model = Qwen3VLForConditionalGeneration.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto"
)

print("Vocabulary:", len(tokenizer))
print("Embedding:", model.get_input_embeddings().weight.shape)

Expected output:

Vocabulary: 151700
Embedding: torch.Size([151700, 2560])

🔤 Testing Newly Added Characters

You can inspect the newly added atoms with:

new_atoms = [
    # Sindhi
    "ٽ", "ڇ", "ڏ", "ڙ", "ڦ", "ڻ",
    # Balochi
    "ۏ",
    # Extended Arabic-Indic digits
    "۰", "۱", "۲", "۳", "۴", "۵", "۶", "۷", "۸", "۹",
    # Symbol
    "؉",
    # Bidi controls
    "\u200d", "\u200f",
    "\u202a", "\u202b", "\u202c", "\u202d", "\u202e",
    # Harakat
    "ٌ", "ٍ", "ٓ", "ٔ", "ٕ", "ٰ",
]

for char in new_atoms:
    token_id = tokenizer.convert_tokens_to_ids(char)
    print(repr(char), "->", token_id)

Expected IDs run from 151669 through 151699 (see the table above).


📈 Token Efficiency

After the surgery, the following text categories tokenize more efficiently:

Category Before After
Eastern Arabic-Indic digits (۰۱۲۳۴۵۶۷۸۹) 2 tokens each 1 token each
Sindhi consonants ٽ ڇ ڏ ڙ ڦ ڻ 2 tokens each 1 token each
Balochi vowel ۏ 2 tokens 1 token
Bidi controls 2 tokens each 1 token each
Select harakat 2 tokens each 1 token each

Approximate efficiency gains on Pashto, Sindhi, Balochi, and Persian/Urdu numeric text: ~10–20% fewer tokens for text that uses these atoms.


⚠️ Status

Tokenizer surgery: ✅ Complete

Vocabulary resize: ✅ Complete

Embedding forensic audit: ✅ Passed

Model save: ✅ Complete

Tokenizer save: ✅ Complete

Hugging Face upload: ✅ Complete

Reload verification: ✅ Complete

Continued pretraining: ⏳ Not performed as part of this surgery

SFT / instruction tuning: ⏳ Not performed as part of this surgery


📜 Forensic Report

A complete forensic report was generated during the surgery:

pashto_urdu_fresh_embedding_forensics_Qwen_Qwen3-VL-4B-Thinking.json

The report contains:

  • tokenizer audit
  • missing-token inventory
  • new token IDs
  • embedding shapes
  • embedding fingerprints
  • duplicate detection
  • pairwise cosine analysis
  • old/new embedding distribution
  • final forensic gate
  • model verification

👤 Author

nassimjp

Pashto AI / Pashto NLP experiments


📄 License

Please refer to the license of the original base model and repository configuration before redistributing modified model weights.


⭐ Project

Qwen3-VL-4B-Thinking-Pashto-Zi

A focused tokenizer surgery on top of an already multilingual base:

Qwen/Qwen3-VL-4B-Thinking
          +
31 new Pashto + Sindhi + Balochi + Arabic-script atoms
          ↓
151,669 → 151,700 vocabulary
          ↓
Multilingual Arabic-script tokenizer foundation
          +
Vision tower preserved intact

The tokenizer now has dedicated vocabulary entries for the missing Pashto, Sindhi, Balochi, and related Arabic-script characters that Qwen3-VL previously split into byte fragments.


---
Downloads last month
35
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nassimjp/Qwen3-VL-4B-Thinking-Pashto-Zi

Finetuned
(30)
this model