🧬 MiniCPM5-1B-Pashto

Pashto + Urdu Tokenizer Surgery Model


📋 Model Overview

This is a tokenizer-modified version of openbmb/MiniCPM5-1B with 46 Pashto and Urdu-specific characters added as single tokens to the vocabulary.

Property Value
Base Model openbmb/MiniCPM5-1B
Original Vocab Size 130,560
New Vocab Size 130,606
Atoms Added 46
Architecture LlamaForCausalLM
Hidden Size 1,536
Layers 24
Context Length 131,072
Embeddings NOT tied (independent input/output)

🔍 Why This Model Exists

MiniCPM5-1B's tokenizer does not natively support many Pashto and Urdu characters as single tokens. This causes:

  • ❌ Characters being split into multiple tokens
  • ❌ Inefficient encoding
  • ❌ Poor representation for Pashto/Urdu text

This model fixes that by surgically adding 46 missing atoms as independent tokens.


📊 Complete Audit & Modification Report

✅ Existing Single-Token Atoms (20)

Atom ID Atom ID
ا 20541 د 57692
ب 79922 ر 37101
پ 120234 ز 124105
ت 55761 س 75848
ح 118797 ش 107172
ع 84439 ک 77297
ف 79312 ل 29673
ق 74411 م 36367
ن 35964 ه 64732
و 41636 ی 46703

⚠️ Split Atoms — FIXED via Surgery (46)

# Atom Old IDs New ID
1 ښ [172, 270] 130560
2 څ [172, 249] 130561
3 ځ [172, 245] 130562
4 ڼ [172, 142] 130563
5 ږ [172, 266] 130564
6 ډ [172, 253] 130565
7 ټ [171, 142] 130566
8 ړ [172, 263] 130567
9 ې [173, 260] 130568
10 ۍ [173, 257] 130569
11 ګ [172, 126] 130570
12 ث [170, 126] 130571
13 ج [170, 127] 130572
14 چ [172, 250] 130573
15 خ [170, 128] 130574
16 ذ [170, 130] 130575
17 ص [170, 135] 130576
18 ض [170, 136] 130577
19 ط [170, 137] 130578
20 ظ [170, 138] 130579
21 غ [170, 140] 130580
22 ئ [170, 121] 130581
23 ے [173, 262] 130582
24 ۀ [173, 244] 130583
25 ٹ [171, 139] 130584
26 ڈ [172, 252] 130585
27 ڑ [172, 261] 130586
28 ں [172, 140] 130587
29 ھ [172, 144] 130588
30 گ [172, 129] 130589
31 أ [170, 118] 130590
32 إ [170, 120] 130591
33 آ [170, 117] 130592
34 ؤ [170, 119] 130593
35 ء [170, 116] 130594
36 ٱ [171, 131] 130595
37 ۰ [173, 130] 130596
38 ۱ [173, 131] 130597
39 ۲ [173, 132] 130598
40 ۳ [173, 133] 130599
41 ۴ [173, 134] 130600
42 ۵ [173, 135] 130601
43 ۶ [173, 136] 130602
44 ۷ [173, 137] 130603
45 ۸ [173, 138] 130604
46 ۹ [173, 139] 130605

🔬 Embedding Forensics

Model Configuration

{
    "architectures": ["LlamaForCausalLM"],
    "vocab_size": 130606,
    "hidden_size": 1536,
    "num_hidden_layers": 24,
    "max_position_embeddings": 131072,
    "tie_word_embeddings": False,
    "initializer_range": 0.02,
}

Input Embedding Statistics (New Atoms)

Metric Value
Mean embedding norm 0.901353
Max cosine similarity 0.25996944

Most Similar Pair (New Atoms)

Atom 1 Atom 2 Cosine Similarity
ح ع 0.25996944

✅ No exact duplicate embeddings detected

Sample New Token Embedding Fingerprints

Atom ID Input Norm Input Mean Input Std
ښ 130560 0.780407 0.000240 0.019918
څ 130561 0.789077 -0.000997 0.020116
ځ 130562 0.762457 0.000163 0.019460
ڼ 130563 0.771574 0.000194 0.019693
ږ 130564 0.795359 -0.000255 0.020299
ډ 130565 0.764082 -0.000497 0.019496
ټ 130566 0.775816 -0.000329 0.019799
ړ 130567 0.748911 0.000773 0.019099
ې 130568 0.784092 -0.000552 0.020005
ۍ 130569 0.779827 0.000641 0.019894

🚀 Usage

Installation

pip install transformers torch accelerate

Basic Loading

from transformers import AutoTokenizer, AutoModelForCausalLM

# Load model and tokenizer
model_id = "nassimjp/MiniCPM5-1B-Pashto"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

Tokenization Check

# Test Pashto characters
pashto_text = "سلام ښه راغلئ"

# Tokenize
tokens = tokenizer.encode(pashto_text, add_special_tokens=False)
print(f"Token IDs: {tokens}")

# Decode back
decoded = tokenizer.decode(tokens)
print(f"Decoded: {decoded}")

# Check individual atoms
for char in "ښڅځڼږډټړېۍګ":
    ids = tokenizer.encode(char, add_special_tokens=False)
    print(f"'{char}' → {ids}")

Text Generation

import torch

# Simple generation
prompt = "ښه راغلئ"
inputs = tokenizer.encode(prompt, return_tensors="pt")

# Generate
with torch.no_grad():
    outputs = model.generate(
        inputs,
        max_new_tokens=50,
        temperature=0.7,
        do_sample=True,
        pad_token_id=tokenizer.eos_token_id
    )

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Batch Processing

texts = [
    "سلام",
    "ښه راغلئ",
    "پښتو ژبه",
    "ستاسو نوم څه دی؟"
]

# Tokenize batch
encodings = tokenizer(
    texts,
    padding=True,
    return_tensors="pt"
)

print(encodings.input_ids.shape)  # (4, sequence_length)

⚙️ Technical Details

Surgery Methodology

  1. Audit Phase

    • Iterated through 66 Pashto/Urdu characters
    • Detected which characters split into multiple tokens
    • Identified 46 characters needing surgery
  2. Token Addition

    • Used tokenizer.add_tokens() to add 46 missing atoms
    • Resized model embeddings with model.resize_token_embeddings()
    • New tokens assigned IDs: 130560 → 130605
  3. Independent Initialization

    • Input embeddings initialized with normal distribution
    • Output embeddings initialized independently
    • Both use initializer_range = 0.02
    • Not tied — maintains separate input and output spaces
  4. Verification

    • Tokenization gate confirmed all 66 atoms as single tokens
    • No exact duplicate embeddings detected
    • Embedding norms and covariances verified

⚠️ Important Notes

Note Description
No Finetuning This model has NOT been finetuned on Pashto/Urdu data
Embeddings Input and output embeddings are NOT tied
New Tokens Added tokens are at the end of vocabulary (IDs 130560+)
Performance Requires finetuning on Pashto/Urdu data for optimal performance
Tokenizer Works out-of-the-box with all Pashto/Urdu atoms

📦 Model Files

MiniCPM5-1B-Pashto/
├── chat_template.jinja             0.01 MB
├── config.json                     0.00 MB
├── generation_config.json          0.00 MB
├── ipashto_surgery_info.txt        0.00 MB
├── model-00001-of-00002.safetensors  1,899.42 MB
├── model-00002-of-00002.safetensors    162.02 MB
├── model.safetensors.index.json    0.02 MB
├── tokenizer.json                  9.44 MB
└── tokenizer_config.json           0.00 MB

Total Size: 2.02 GB

🔗 Related Resources


📝 Citation

If you use this model, please cite the original MiniCPM paper:

@article{minicpm2024,
  title={MiniCPM: Unveiling the Potential of Small Language Models},
  author={Hu, Shengding and Ding, Ning and others},
  journal={arXiv preprint arXiv:2404.06395},
  year={2024}
}

📄 License

This model is released under the Apache License 2.0.


🤝 Acknowledgements

  • OpenBMB for developing and releasing MiniCPM5-1B
  • Hugging Face for the transformers library and model hosting
  • Kaggle for providing the compute environment

🧪 Testing Results

All 66 Pashto/Urdu atoms successfully tokenize as single tokens:

Category Count Status
Pashto-Specific 11 ✅ All single tokens
Shared Arabic/Persian 30 ✅ All single tokens
Positional Forms 3 ✅ All single tokens
Urdu/South Asian 6 ✅ All single tokens
Arabic Variants 6 ✅ All single tokens
Eastern Arabic Digits 10 ✅ All single tokens
TOTAL 66 ✅ 100% Coverage

📞 Contact & Support


Created with ❤️ for the Pashto and Urdu communities


---
Downloads last month
1,369
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nassimjp/MiniCPM5-1B-Pashto

Finetuned
(57)
this model
Quantizations
1 model

Paper for nassimjp/MiniCPM5-1B-Pashto