Instructions to use nassimjp/MiniCPM5-1B-Pashto with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nassimjp/MiniCPM5-1B-Pashto with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nassimjp/MiniCPM5-1B-Pashto") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nassimjp/MiniCPM5-1B-Pashto") model = AutoModelForCausalLM.from_pretrained("nassimjp/MiniCPM5-1B-Pashto", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nassimjp/MiniCPM5-1B-Pashto with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nassimjp/MiniCPM5-1B-Pashto" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nassimjp/MiniCPM5-1B-Pashto", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nassimjp/MiniCPM5-1B-Pashto
- SGLang
How to use nassimjp/MiniCPM5-1B-Pashto with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nassimjp/MiniCPM5-1B-Pashto" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nassimjp/MiniCPM5-1B-Pashto", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nassimjp/MiniCPM5-1B-Pashto" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nassimjp/MiniCPM5-1B-Pashto", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use nassimjp/MiniCPM5-1B-Pashto with Docker Model Runner:
docker model run hf.co/nassimjp/MiniCPM5-1B-Pashto
- 🧬 MiniCPM5-1B-Pashto
🧬 MiniCPM5-1B-Pashto
Pashto + Urdu Tokenizer Surgery Model
📋 Model Overview
This is a tokenizer-modified version of openbmb/MiniCPM5-1B with 46 Pashto and Urdu-specific characters added as single tokens to the vocabulary.
| Property | Value |
|---|---|
| Base Model | openbmb/MiniCPM5-1B |
| Original Vocab Size | 130,560 |
| New Vocab Size | 130,606 |
| Atoms Added | 46 |
| Architecture | LlamaForCausalLM |
| Hidden Size | 1,536 |
| Layers | 24 |
| Context Length | 131,072 |
| Embeddings | NOT tied (independent input/output) |
🔍 Why This Model Exists
MiniCPM5-1B's tokenizer does not natively support many Pashto and Urdu characters as single tokens. This causes:
- ❌ Characters being split into multiple tokens
- ❌ Inefficient encoding
- ❌ Poor representation for Pashto/Urdu text
This model fixes that by surgically adding 46 missing atoms as independent tokens.
📊 Complete Audit & Modification Report
✅ Existing Single-Token Atoms (20)
| Atom | ID | Atom | ID |
|---|---|---|---|
ا |
20541 | د |
57692 |
ب |
79922 | ر |
37101 |
پ |
120234 | ز |
124105 |
ت |
55761 | س |
75848 |
ح |
118797 | ش |
107172 |
ع |
84439 | ک |
77297 |
ف |
79312 | ل |
29673 |
ق |
74411 | م |
36367 |
ن |
35964 | ه |
64732 |
و |
41636 | ی |
46703 |
⚠️ Split Atoms — FIXED via Surgery (46)
| # | Atom | Old IDs | New ID |
|---|---|---|---|
| 1 | ښ |
[172, 270] | 130560 |
| 2 | څ |
[172, 249] | 130561 |
| 3 | ځ |
[172, 245] | 130562 |
| 4 | ڼ |
[172, 142] | 130563 |
| 5 | ږ |
[172, 266] | 130564 |
| 6 | ډ |
[172, 253] | 130565 |
| 7 | ټ |
[171, 142] | 130566 |
| 8 | ړ |
[172, 263] | 130567 |
| 9 | ې |
[173, 260] | 130568 |
| 10 | ۍ |
[173, 257] | 130569 |
| 11 | ګ |
[172, 126] | 130570 |
| 12 | ث |
[170, 126] | 130571 |
| 13 | ج |
[170, 127] | 130572 |
| 14 | چ |
[172, 250] | 130573 |
| 15 | خ |
[170, 128] | 130574 |
| 16 | ذ |
[170, 130] | 130575 |
| 17 | ص |
[170, 135] | 130576 |
| 18 | ض |
[170, 136] | 130577 |
| 19 | ط |
[170, 137] | 130578 |
| 20 | ظ |
[170, 138] | 130579 |
| 21 | غ |
[170, 140] | 130580 |
| 22 | ئ |
[170, 121] | 130581 |
| 23 | ے |
[173, 262] | 130582 |
| 24 | ۀ |
[173, 244] | 130583 |
| 25 | ٹ |
[171, 139] | 130584 |
| 26 | ڈ |
[172, 252] | 130585 |
| 27 | ڑ |
[172, 261] | 130586 |
| 28 | ں |
[172, 140] | 130587 |
| 29 | ھ |
[172, 144] | 130588 |
| 30 | گ |
[172, 129] | 130589 |
| 31 | أ |
[170, 118] | 130590 |
| 32 | إ |
[170, 120] | 130591 |
| 33 | آ |
[170, 117] | 130592 |
| 34 | ؤ |
[170, 119] | 130593 |
| 35 | ء |
[170, 116] | 130594 |
| 36 | ٱ |
[171, 131] | 130595 |
| 37 | ۰ |
[173, 130] | 130596 |
| 38 | ۱ |
[173, 131] | 130597 |
| 39 | ۲ |
[173, 132] | 130598 |
| 40 | ۳ |
[173, 133] | 130599 |
| 41 | ۴ |
[173, 134] | 130600 |
| 42 | ۵ |
[173, 135] | 130601 |
| 43 | ۶ |
[173, 136] | 130602 |
| 44 | ۷ |
[173, 137] | 130603 |
| 45 | ۸ |
[173, 138] | 130604 |
| 46 | ۹ |
[173, 139] | 130605 |
🔬 Embedding Forensics
Model Configuration
{
"architectures": ["LlamaForCausalLM"],
"vocab_size": 130606,
"hidden_size": 1536,
"num_hidden_layers": 24,
"max_position_embeddings": 131072,
"tie_word_embeddings": False,
"initializer_range": 0.02,
}
Input Embedding Statistics (New Atoms)
| Metric | Value |
|---|---|
| Mean embedding norm | 0.901353 |
| Max cosine similarity | 0.25996944 |
Most Similar Pair (New Atoms)
| Atom 1 | Atom 2 | Cosine Similarity |
|---|---|---|
ح |
ع |
0.25996944 |
✅ No exact duplicate embeddings detected
Sample New Token Embedding Fingerprints
| Atom | ID | Input Norm | Input Mean | Input Std |
|---|---|---|---|---|
ښ |
130560 | 0.780407 | 0.000240 | 0.019918 |
څ |
130561 | 0.789077 | -0.000997 | 0.020116 |
ځ |
130562 | 0.762457 | 0.000163 | 0.019460 |
ڼ |
130563 | 0.771574 | 0.000194 | 0.019693 |
ږ |
130564 | 0.795359 | -0.000255 | 0.020299 |
ډ |
130565 | 0.764082 | -0.000497 | 0.019496 |
ټ |
130566 | 0.775816 | -0.000329 | 0.019799 |
ړ |
130567 | 0.748911 | 0.000773 | 0.019099 |
ې |
130568 | 0.784092 | -0.000552 | 0.020005 |
ۍ |
130569 | 0.779827 | 0.000641 | 0.019894 |
🚀 Usage
Installation
pip install transformers torch accelerate
Basic Loading
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load model and tokenizer
model_id = "nassimjp/MiniCPM5-1B-Pashto"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
Tokenization Check
# Test Pashto characters
pashto_text = "سلام ښه راغلئ"
# Tokenize
tokens = tokenizer.encode(pashto_text, add_special_tokens=False)
print(f"Token IDs: {tokens}")
# Decode back
decoded = tokenizer.decode(tokens)
print(f"Decoded: {decoded}")
# Check individual atoms
for char in "ښڅځڼږډټړېۍګ":
ids = tokenizer.encode(char, add_special_tokens=False)
print(f"'{char}' → {ids}")
Text Generation
import torch
# Simple generation
prompt = "ښه راغلئ"
inputs = tokenizer.encode(prompt, return_tensors="pt")
# Generate
with torch.no_grad():
outputs = model.generate(
inputs,
max_new_tokens=50,
temperature=0.7,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Batch Processing
texts = [
"سلام",
"ښه راغلئ",
"پښتو ژبه",
"ستاسو نوم څه دی؟"
]
# Tokenize batch
encodings = tokenizer(
texts,
padding=True,
return_tensors="pt"
)
print(encodings.input_ids.shape) # (4, sequence_length)
⚙️ Technical Details
Surgery Methodology
Audit Phase
- Iterated through 66 Pashto/Urdu characters
- Detected which characters split into multiple tokens
- Identified 46 characters needing surgery
Token Addition
- Used
tokenizer.add_tokens()to add 46 missing atoms - Resized model embeddings with
model.resize_token_embeddings() - New tokens assigned IDs: 130560 → 130605
- Used
Independent Initialization
- Input embeddings initialized with normal distribution
- Output embeddings initialized independently
- Both use
initializer_range = 0.02 - Not tied — maintains separate input and output spaces
Verification
- Tokenization gate confirmed all 66 atoms as single tokens
- No exact duplicate embeddings detected
- Embedding norms and covariances verified
⚠️ Important Notes
| Note | Description |
|---|---|
| No Finetuning | This model has NOT been finetuned on Pashto/Urdu data |
| Embeddings | Input and output embeddings are NOT tied |
| New Tokens | Added tokens are at the end of vocabulary (IDs 130560+) |
| Performance | Requires finetuning on Pashto/Urdu data for optimal performance |
| Tokenizer | Works out-of-the-box with all Pashto/Urdu atoms |
📦 Model Files
MiniCPM5-1B-Pashto/
├── chat_template.jinja 0.01 MB
├── config.json 0.00 MB
├── generation_config.json 0.00 MB
├── ipashto_surgery_info.txt 0.00 MB
├── model-00001-of-00002.safetensors 1,899.42 MB
├── model-00002-of-00002.safetensors 162.02 MB
├── model.safetensors.index.json 0.02 MB
├── tokenizer.json 9.44 MB
└── tokenizer_config.json 0.00 MB
Total Size: 2.02 GB
🔗 Related Resources
- Base Model: openbmb/MiniCPM5-1B
- Original Paper: MiniCPM: Unveiling the Potential of Small Language Models
- GitHub: OpenBMB/MiniCPM
📝 Citation
If you use this model, please cite the original MiniCPM paper:
@article{minicpm2024,
title={MiniCPM: Unveiling the Potential of Small Language Models},
author={Hu, Shengding and Ding, Ning and others},
journal={arXiv preprint arXiv:2404.06395},
year={2024}
}
📄 License
This model is released under the Apache License 2.0.
🤝 Acknowledgements
- OpenBMB for developing and releasing MiniCPM5-1B
- Hugging Face for the transformers library and model hosting
- Kaggle for providing the compute environment
🧪 Testing Results
All 66 Pashto/Urdu atoms successfully tokenize as single tokens:
| Category | Count | Status |
|---|---|---|
| Pashto-Specific | 11 | ✅ All single tokens |
| Shared Arabic/Persian | 30 | ✅ All single tokens |
| Positional Forms | 3 | ✅ All single tokens |
| Urdu/South Asian | 6 | ✅ All single tokens |
| Arabic Variants | 6 | ✅ All single tokens |
| Eastern Arabic Digits | 10 | ✅ All single tokens |
| TOTAL | 66 | ✅ 100% Coverage |
📞 Contact & Support
- Model Page: nassimjp/MiniCPM5-1B-Pashto
- Issues: Please open an issue on the model page
- Suggestions: Feedback welcome!
Created with ❤️ for the Pashto and Urdu communities
---
- Downloads last month
- 1,369