Qwen2.5-0.5B-Arabic-Classic (Pruned)
Base model: Qwen/Qwen2.5-0.5B
Model type: Causal Language Model (Base)
Vocab size: 30,557 tokens (β80% from 151,665)
Parameters: ~385M (after vocab expansion)
Context length: 1,024 tokens (training) / 32,768 tokens (architecture)
License: Apache 2.0
π Summary
This model is the result of Continued Pre-Training (CPT) of Qwen2.5-0.5B on ~300 million tokens of classical Arabic text from the Shamela/Waqfeya collection of Islamic heritage books (kitab turath). The model is specifically optimized for classical Arabic, shar'i, and fiqh domains through two key innovations:
- Leaf-Based Vocabulary Pruning β Reduces vocabulary from 151,665 to 30,557 tokens (80% pruning) while preserving a valid BPE merge structure with zero unreachable tokens, based on the paper "Teaching Old Tokenizers New Words" (Purason et al., 2026).
- GaLore Optimizer β Full-parameter training with memory-efficient low-rank projection for attention & MLP matrices, enabling training on 2Γ Tesla T4 16GB GPUs.
ποΈ Architecture & Modifications
| Component | Specification |
|---|---|
| Architecture | Transformer decoder with RoPE, SwiGLU, RMSNorm, GQA |
| Layers | 24 |
| Hidden size | 896 |
| Attention heads | 14 (Q) / 2 (KV) |
| Original vocab | 151,665 |
| Pruned vocab | 30,556 |
| Expanded vocab | 30,557 (+1 special token </SEP>) |
| Embedding shape | (30,557, 896) |
| Tied embeddings | Yes (embed_tokens β lm_head) |
Vocabulary Pruning
- Method: Leaf-based frequency pruning β only leaf tokens (not used as input to any other merge) are removed based on low frequency in the target corpus.
- Protected tokens: 278 atomic tokens + special tokens.
- Tokens pruned: 121,109 (79.9%).
- New unreachable tokens: 0 (verified).
Additional Special Token
</SEP>(ID: 30556) β Used as a document separator during packing.
π§ͺ Training Details
Dataset
| Source | Type | Estimated Tokens |
|---|---|---|
| Shamela/Waqfeya | Classical Arabic books (tafsir, hadith, fiqh, aqidah, nahwu) | ~300 million tokens |
| Parquet shards | Pre-packed to 1,024 sequence length | 26,902 steps/epoch |
Hyperparameters
| Parameter | Value |
|---|---|
| Batch size per GPU | 16 |
| Gradient accumulation | 4 |
| Global batch (2Γ T4) | 128 |
| Learning rate (GaLore) | 1e-4 |
| Learning rate (embed/lm_head) | 5e-4 |
| LR scheduler | Cosine with warmup (200 steps) |
| Weight decay | 0.01 |
| GaLore rank | 128 |
| GaLore update gap | 400 |
| GaLore scale | 0.25 |
| Precision | FP16 (T4) |
| Gradient checkpointing | Enabled |
| Optimizer | GaLoreAdamW (3 parameter groups) |
Hardware
- 2Γ NVIDIA Tesla T4 (16 GB VRAM each)
- PyTorch DDP via
torchrun --nproc_per_node=2 - Training time: ~35 seconds/step (including eval & save every 100 steps)
π Evaluation
Perplexity & Bits-per-Byte (Held-out Test, n=50)
| Metric | Base Model (Qwen2.5-0.5B) | This Model | Improvement |
|---|---|---|---|
| Avg Loss | 2.359 | 1.655 | β29.8% |
| Perplexity | 10.58 | 5.23 | β50.6% |
| Bits/Byte | 0.698 | 0.490 | β29.8% |
Per-Genre Evaluation (Domain-specific)
| Genre | Base PPL | Our PPL | Improvement |
|---|---|---|---|
| Quran | 3.36 | 1.69 | β49.6% |
| Hadith | 2.63 | 1.40 | β47.0% |
| Tafsir | 4.34 | 2.45 | β43.7% |
| Fiqh (Classical) | 5.72 | 2.41 | β57.8% |
| Prose Turath | 5.21 | 3.02 | β42.0% |
All differences are statistically significant (bootstrap CI 95%, p < 0.05).
Tashkeel Robustness
- This model: 0.154 (lower = more robust to harakat variations)
- Base: 0.249
Generation Speed
- This model: 9.89 tokens/sec
- Base: 8.37 tokens/sec (+18% faster, due to smaller vocabulary)
Diversity (Distinct-n)
| Model | Distinct-1 | Distinct-2 | Distinct-3 |
|---|---|---|---|
| This model | 0.878 | 0.989 | 1.0 |
| Base | 0.880 | 0.957 | 1.0 |
π‘ Generation Examples
Hadith Continuation
Input: ΨΩΨ―ΩΩΨ«ΩΩΩΨ§ Ψ³ΩΩΩΩΩΨ§ΩΩ ΨΉΩΩΩ Ψ§ΩΨ²ΩΩΩΩΨ±ΩΩΩΩ
Output: ΨΩΨ―ΩΩΨ«ΩΩΩΨ§ Ψ³ΩΩΩΩΩΨ§ΩΩ ΨΉΩΩΩ Ψ§ΩΨ²ΩΩΩΩΨ±ΩΩΩΩ ΨΉΩΩΩ ΨΉΩΨ¨ΩΩΩΨ―Ω Ψ§ΩΩΩΩΩΩ Ψ¨ΩΩΩ ΨΉΩΨ¨ΩΨ―Ω Ψ§ΩΩΩΩΩΩ Ψ¨ΩΩΩ ΨΉΩΨͺΩΨ¨ΩΨ©Ω ΨΉΩΩΩ Ψ£ΩΨ¨ΩΩΩΩ ΩΩΨ§ΩΩ : ΩΩΨ§ΩΩ Ψ±...
Fiqh Ruling
Input: ΩΩΨ§ΩΨ―ΩΩΩΩΩΩΩ ΨΉΩΩΩΩ ΩΩΨ¬ΩΩΨ¨Ω ΩΩΨ°ΩΩΩ Ψ§ΩΩΩ ΩΨ³ΩΨ£ΩΩΩΨ©Ω ΨΉΩΩΩΨ―Ω Ψ§ΩΩΩΩΩΩΩΩΨ§Ψ‘Ω
Output: ΩΩΨ§ΩΨ―ΩΩΩΩΩΩΩ ΨΉΩΩΩΩ ΩΩΨ¬ΩΩΨ¨Ω ΩΩΨ°ΩΩΩ Ψ§ΩΩΩ ΩΨ³ΩΨ£ΩΩΩΨ©Ω ΨΉΩΩΩΨ―Ω Ψ§ΩΩΩΩΩΩΩΩΨ§Ψ‘Ω Ψ£ΩΩΩΩΩΩΨ§ Ψ₯ΩΨ°ΩΨ§ ΩΩΨ§ΩΩΨͺΩ Ω ΩΨΩΨ±ΩΩΩ ΩΨ©Ω ΩΩΨ₯ΩΩΩΩΩΩΨ§ ΨͺΩΩΩΩΩΩ Ω ΩΨΩΨ±ΩΩΩ ΩΨ©Ω Ψ¨ΩΨ§ΩΩ...
Tafsir Style
Input: ΩΩΩΩΩΩΩΩ ΨͺΩΨΉΩΨ§ΩΩΩ ΩΩΨ§Ψ΅ΩΨ¨ΩΨ±Ω ΩΩΩΩΨ³ΩΩΩ Ω ΩΨΉΩ Ψ§ΩΩΩΨ°ΩΩΩΩ
Output: ΩΩΩΩΩΩΩΩ ΨͺΩΨΉΩΨ§ΩΩΩ ΩΩΨ§Ψ΅ΩΨ¨ΩΨ±Ω ΩΩΩΩΨ³ΩΩΩ Ω ΩΨΉΩ Ψ§ΩΩΩΨ°ΩΩΩΩ ΩΩΨΉΩΩΩΩ Ψ¨ΩΨ£ΩΨ±ΩΨΩΨ§Ω ΩΩΩΩ Ω Ψ£ΩΩΩΩΨ³ΩΩΩΩ Ω ΩΩΨ£ΩΨ²ΩΩΩΨ¬ΩΩΩΩ Ω ΩΩΨ°ΩΨ±ΩΩΩΩΩΨͺΩΩΩΩ Ω Ψ₯ΩΩΩΩ...
π Usage
Requirements
pip install transformers>=4.37.0 torch
Load Model
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Ik45/qwen2.5-0.5b-arabic-classic-pruned"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
model_name,
trust_remote_code=True
)
# Generate
prompt = "ΩΩΨ§ΩΩ Ψ§ΩΨ₯Ω
Ψ§Ω
Ω
Ψ§ΩΩ Ψ±ΨΩ
Ω Ψ§ΩΩΩ:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=True,
temperature=0.7,
top_p=0.9
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
β οΈ Important Notes
- Base Model β This is a base model (not instruct/chat). Not recommended for direct dialogue. Use SFT/DPO first if you want to use it as a chatbot.
- Domain Specific β The model is highly specialized for classical Arabic text. Performance on Modern Standard Arabic or other languages may degrade due to aggressive vocabulary pruning.
- Context Length β Trained on 1,024 tokens, but the architecture supports up to 32,768 tokens. For extremely long contexts, additional fine-tuning with RoPE scaling is recommended.
- Vocabulary Pruning β Since 80% of the original Qwen vocabulary was removed, tokenization of non-Arabic text will produce longer sequences (fallback to subword/character-level).
π References
Papers & Techniques
- Leaf-Based Vocabulary Pruning: Purason et al., "Teaching Old Tokenizers New Words", arXiv:2512.03989v2 (2026).
- GaLore: Zhao et al., "GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection", arXiv:2403.03507.
- Qwen2.5: Qwen Team, "Qwen2.5: A Party of Foundation Models", 2024.
Dataset
- Shamela/Waqfeya β Digital collection of classical Islamic books.
- Pretraining dataset: ~300 million tokens from various genres (Quran, Hadith, Tafsir, Fiqh, Aqidah, Nahwu, Balagha).
π Acknowledgments
- Qwen Team for the Qwen2.5-0.5B base model.
- Purason et al. for the leaf-based vocabulary pruning method.
- Shamela/Waqfeya for the classical Arabic book dataset.
π License
This model is released under the Apache 2.0 license, same as the base Qwen2.5-0.5B model. """
- Downloads last month
- 15