Qwen2.5-0.5B-Arabic-Classic (Pruned)

Base model: Qwen/Qwen2.5-0.5B
Model type: Causal Language Model (Base)
Vocab size: 30,557 tokens (↓80% from 151,665)
Parameters: ~385M (after vocab expansion)
Context length: 1,024 tokens (training) / 32,768 tokens (architecture)
License: Apache 2.0


πŸ“‹ Summary

This model is the result of Continued Pre-Training (CPT) of Qwen2.5-0.5B on ~300 million tokens of classical Arabic text from the Shamela/Waqfeya collection of Islamic heritage books (kitab turath). The model is specifically optimized for classical Arabic, shar'i, and fiqh domains through two key innovations:

  1. Leaf-Based Vocabulary Pruning β€” Reduces vocabulary from 151,665 to 30,557 tokens (80% pruning) while preserving a valid BPE merge structure with zero unreachable tokens, based on the paper "Teaching Old Tokenizers New Words" (Purason et al., 2026).
  2. GaLore Optimizer β€” Full-parameter training with memory-efficient low-rank projection for attention & MLP matrices, enabling training on 2Γ— Tesla T4 16GB GPUs.

πŸ—οΈ Architecture & Modifications

Component Specification
Architecture Transformer decoder with RoPE, SwiGLU, RMSNorm, GQA
Layers 24
Hidden size 896
Attention heads 14 (Q) / 2 (KV)
Original vocab 151,665
Pruned vocab 30,556
Expanded vocab 30,557 (+1 special token </SEP>)
Embedding shape (30,557, 896)
Tied embeddings Yes (embed_tokens ↔ lm_head)

Vocabulary Pruning

  • Method: Leaf-based frequency pruning β€” only leaf tokens (not used as input to any other merge) are removed based on low frequency in the target corpus.
  • Protected tokens: 278 atomic tokens + special tokens.
  • Tokens pruned: 121,109 (79.9%).
  • New unreachable tokens: 0 (verified).

Additional Special Token

  • </SEP> (ID: 30556) β€” Used as a document separator during packing.

πŸ§ͺ Training Details

Dataset

Source Type Estimated Tokens
Shamela/Waqfeya Classical Arabic books (tafsir, hadith, fiqh, aqidah, nahwu) ~300 million tokens
Parquet shards Pre-packed to 1,024 sequence length 26,902 steps/epoch

Hyperparameters

Parameter Value
Batch size per GPU 16
Gradient accumulation 4
Global batch (2Γ— T4) 128
Learning rate (GaLore) 1e-4
Learning rate (embed/lm_head) 5e-4
LR scheduler Cosine with warmup (200 steps)
Weight decay 0.01
GaLore rank 128
GaLore update gap 400
GaLore scale 0.25
Precision FP16 (T4)
Gradient checkpointing Enabled
Optimizer GaLoreAdamW (3 parameter groups)

Hardware

  • 2Γ— NVIDIA Tesla T4 (16 GB VRAM each)
  • PyTorch DDP via torchrun --nproc_per_node=2
  • Training time: ~35 seconds/step (including eval & save every 100 steps)

πŸ“Š Evaluation

Perplexity & Bits-per-Byte (Held-out Test, n=50)

Metric Base Model (Qwen2.5-0.5B) This Model Improvement
Avg Loss 2.359 1.655 ↓29.8%
Perplexity 10.58 5.23 ↓50.6%
Bits/Byte 0.698 0.490 ↓29.8%

Per-Genre Evaluation (Domain-specific)

Genre Base PPL Our PPL Improvement
Quran 3.36 1.69 ↓49.6%
Hadith 2.63 1.40 ↓47.0%
Tafsir 4.34 2.45 ↓43.7%
Fiqh (Classical) 5.72 2.41 ↓57.8%
Prose Turath 5.21 3.02 ↓42.0%

All differences are statistically significant (bootstrap CI 95%, p < 0.05).

Tashkeel Robustness

  • This model: 0.154 (lower = more robust to harakat variations)
  • Base: 0.249

Generation Speed

  • This model: 9.89 tokens/sec
  • Base: 8.37 tokens/sec (+18% faster, due to smaller vocabulary)

Diversity (Distinct-n)

Model Distinct-1 Distinct-2 Distinct-3
This model 0.878 0.989 1.0
Base 0.880 0.957 1.0

πŸ’‘ Generation Examples

Hadith Continuation

Input: Ψ­ΩŽΨ―ΩŽΩ‘Ψ«ΩŽΩ†ΩŽΨ§ Ψ³ΩΩΩ’ΩŠΩŽΨ§Ω†Ω ΨΉΩŽΩ†Ω Ψ§Ω„Ψ²ΩΩ‘Ω‡Ω’Ψ±ΩΩŠΩΩ‘
Output: Ψ­ΩŽΨ―ΩŽΩ‘Ψ«ΩŽΩ†ΩŽΨ§ Ψ³ΩΩΩ’ΩŠΩŽΨ§Ω†Ω ΨΉΩŽΩ†Ω Ψ§Ω„Ψ²ΩΩ‘Ω‡Ω’Ψ±ΩΩŠΩΩ‘ ΨΉΩŽΩ†Ω’ ΨΉΩΨ¨ΩŽΩŠΩ’Ψ―Ω Ψ§Ω„Ω„ΩŽΩ‘Ω‡Ω بْنِ ΨΉΩŽΨ¨Ω’Ψ―Ω Ψ§Ω„Ω„ΩŽΩ‘Ω‡Ω بْنِ عُΨͺΩ’Ψ¨ΩŽΨ©ΩŽ ΨΉΩŽΩ†Ω’ Ψ£ΩŽΨ¨ΩΩŠΩ‡Ω Ω‚ΩŽΨ§Ω„ΩŽ : ΩƒΩŽΨ§Ω†ΩŽ Ψ±...

Fiqh Ruling

Input: ΩˆΩŽΨ§Ω„Ψ―ΩŽΩ‘Ω„ΩΩŠΩ„Ω ΨΉΩŽΩ„ΩŽΩ‰ وُجُوبِ Ω‡ΩŽΨ°ΩΩ‡Ω Ψ§Ω„Ω’Ω…ΩŽΨ³Ω’Ψ£ΩŽΩ„ΩŽΨ©Ω ΨΉΩΩ†Ω’Ψ―ΩŽ Ψ§Ω„Ω’ΩΩΩ‚ΩŽΩ‡ΩŽΨ§Ψ‘Ω
Output: ΩˆΩŽΨ§Ω„Ψ―ΩŽΩ‘Ω„ΩΩŠΩ„Ω ΨΉΩŽΩ„ΩŽΩ‰ وُجُوبِ Ω‡ΩŽΨ°ΩΩ‡Ω Ψ§Ω„Ω’Ω…ΩŽΨ³Ω’Ψ£ΩŽΩ„ΩŽΨ©Ω ΨΉΩΩ†Ω’Ψ―ΩŽ Ψ§Ω„Ω’ΩΩΩ‚ΩŽΩ‡ΩŽΨ§Ψ‘Ω Ψ£ΩŽΩ†ΩŽΩ‘Ω‡ΩŽΨ§ Ψ₯ِذَا ΩƒΩŽΨ§Ω†ΩŽΨͺΩ’ Ω…ΩΨ­ΩŽΨ±ΩŽΩ‘Ω…ΩŽΨ©ΩŒ فَΨ₯ΩΩ†ΩŽΩ‘Ω‡ΩŽΨ§ ΨͺΩŽΩƒΩΩˆΩ†Ω Ω…ΩΨ­ΩŽΨ±ΩŽΩ‘Ω…ΩŽΨ©ΩŒ بِالْ...

Tafsir Style

Input: Ω‚ΩŽΩˆΩ’Ω„ΩΩ‡Ω ΨͺΩŽΨΉΩŽΨ§Ω„ΩŽΩ‰ ΩˆΩŽΨ§Ψ΅Ω’Ψ¨ΩΨ±Ω’ Ω†ΩŽΩΩ’Ψ³ΩŽΩƒΩŽ Ω…ΩŽΨΉΩŽ Ψ§Ω„ΩŽΩ‘Ψ°ΩΩŠΩ†ΩŽ
Output: Ω‚ΩŽΩˆΩ’Ω„ΩΩ‡Ω ΨͺΩŽΨΉΩŽΨ§Ω„ΩŽΩ‰ ΩˆΩŽΨ§Ψ΅Ω’Ψ¨ΩΨ±Ω’ Ω†ΩŽΩΩ’Ψ³ΩŽΩƒΩŽ Ω…ΩŽΨΉΩŽ Ψ§Ω„ΩŽΩ‘Ψ°ΩΩŠΩ†ΩŽ ΩŠΩŽΨΉΩ’Ω†ΩΩŠ Ψ¨ΩΨ£ΩŽΨ±Ω’Ψ­ΩŽΨ§Ω…ΩΩ‡ΩΩ…Ω’ Ψ£ΩŽΩ†ΩΩΨ³ΩŽΩ‡ΩΩ…Ω’ ΩˆΩŽΨ£ΩŽΨ²Ω’ΩˆΩŽΨ¬ΩŽΩ‡ΩΩ…Ω’ ΩˆΩŽΨ°ΩΨ±ΩΩ‘ΩŠΩŽΩ‘ΨͺΩŽΩ‡ΩΩ…Ω’ Ψ₯ΩΩ†ΩŽΩ‘...


πŸš€ Usage

Requirements

pip install transformers>=4.37.0 torch

Load Model

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Ik45/qwen2.5-0.5b-arabic-classic-pruned"

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
    model_name,
    trust_remote_code=True
)

# Generate
prompt = "Ω‚ΩŽΨ§Ω„ΩŽ Ψ§Ω„Ψ₯Ω…Ψ§Ω… Ω…Ψ§Ω„Ωƒ Ψ±Ψ­Ω…Ω‡ Ψ§Ω„Ω„Ω‡:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=128,
    do_sample=True,
    temperature=0.7,
    top_p=0.9
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

⚠️ Important Notes

  1. Base Model β€” This is a base model (not instruct/chat). Not recommended for direct dialogue. Use SFT/DPO first if you want to use it as a chatbot.
  2. Domain Specific β€” The model is highly specialized for classical Arabic text. Performance on Modern Standard Arabic or other languages may degrade due to aggressive vocabulary pruning.
  3. Context Length β€” Trained on 1,024 tokens, but the architecture supports up to 32,768 tokens. For extremely long contexts, additional fine-tuning with RoPE scaling is recommended.
  4. Vocabulary Pruning β€” Since 80% of the original Qwen vocabulary was removed, tokenization of non-Arabic text will produce longer sequences (fallback to subword/character-level).

πŸ“š References

Papers & Techniques

  • Leaf-Based Vocabulary Pruning: Purason et al., "Teaching Old Tokenizers New Words", arXiv:2512.03989v2 (2026).
  • GaLore: Zhao et al., "GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection", arXiv:2403.03507.
  • Qwen2.5: Qwen Team, "Qwen2.5: A Party of Foundation Models", 2024.

Dataset

  • Shamela/Waqfeya β€” Digital collection of classical Islamic books.
  • Pretraining dataset: ~300 million tokens from various genres (Quran, Hadith, Tafsir, Fiqh, Aqidah, Nahwu, Balagha).

πŸ™ Acknowledgments


πŸ“„ License

This model is released under the Apache 2.0 license, same as the base Qwen2.5-0.5B model. """

Downloads last month
15
Safetensors
Model size
0.4B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Papers for Ik45/qwen2.5-0.5b-arabic-classic