--- language: ar tags: - qwen - qwen2.5 - arabic - classical-arabic - islamic-texts - continued-pretraining - vocabulary-pruning - galore license: apache-2.0 datasets: - shamela-waqfeya - turath-classical-texts --- # Qwen2.5-0.5B-Arabic-Classic (Pruned) **Base model:** [Qwen/Qwen2.5-0.5B](https://huggingface.co/Qwen/Qwen2.5-0.5B) **Model type:** Causal Language Model (Base) **Vocab size:** 30,557 tokens (โ†“80% from 151,665) **Parameters:** ~385M (after vocab expansion) **Context length:** 1,024 tokens (training) / 32,768 tokens (architecture) **License:** Apache 2.0 --- ## ๐Ÿ“‹ Summary This model is the result of **Continued Pre-Training (CPT)** of [Qwen2.5-0.5B](https://huggingface.co/Qwen/Qwen2.5-0.5B) on **~300 million tokens** of classical Arabic text from the Shamela/Waqfeya collection of Islamic heritage books (kitab turath). The model is specifically optimized for **classical Arabic, shar'i, and fiqh** domains through two key innovations: 1. **Leaf-Based Vocabulary Pruning** โ€” Reduces vocabulary from 151,665 to 30,557 tokens (80% pruning) while preserving a valid BPE merge structure with zero unreachable tokens, based on the paper *"Teaching Old Tokenizers New Words"* (Purason et al., 2026). 2. **GaLore Optimizer** โ€” Full-parameter training with memory-efficient low-rank projection for attention & MLP matrices, enabling training on 2ร— Tesla T4 16GB GPUs. --- ## ๐Ÿ—๏ธ Architecture & Modifications | Component | Specification | |-----------|---------------| | Architecture | Transformer decoder with RoPE, SwiGLU, RMSNorm, GQA | | Layers | 24 | | Hidden size | 896 | | Attention heads | 14 (Q) / 2 (KV) | | Original vocab | 151,665 | | **Pruned vocab** | **30,556** | | **Expanded vocab** | **30,557** (+1 special token ``) | | Embedding shape | (30,557, 896) | | Tied embeddings | Yes (embed_tokens โ†” lm_head) | ### Vocabulary Pruning - **Method:** Leaf-based frequency pruning โ€” only *leaf* tokens (not used as input to any other merge) are removed based on low frequency in the target corpus. - **Protected tokens:** 278 atomic tokens + special tokens. - **Tokens pruned:** 121,109 (79.9%). - **New unreachable tokens:** 0 (verified). ### Additional Special Token - `` (ID: 30556) โ€” Used as a document separator during packing. --- ## ๐Ÿงช Training Details ### Dataset | Source | Type | Estimated Tokens | |--------|------|------------------| | Shamela/Waqfeya | Classical Arabic books (tafsir, hadith, fiqh, aqidah, nahwu) | ~300 million tokens | | Parquet shards | Pre-packed to 1,024 sequence length | 26,902 steps/epoch | ### Hyperparameters | Parameter | Value | |-----------|-------| | Batch size per GPU | 16 | | Gradient accumulation | 4 | | Global batch (2ร— T4) | 128 | | Learning rate (GaLore) | 1e-4 | | Learning rate (embed/lm_head) | 5e-4 | | LR scheduler | Cosine with warmup (200 steps) | | Weight decay | 0.01 | | GaLore rank | 128 | | GaLore update gap | 400 | | GaLore scale | 0.25 | | Precision | FP16 (T4) | | Gradient checkpointing | Enabled | | Optimizer | GaLoreAdamW (3 parameter groups) | ### Hardware - **2ร— NVIDIA Tesla T4** (16 GB VRAM each) - **PyTorch DDP** via `torchrun --nproc_per_node=2` - Training time: ~35 seconds/step (including eval & save every 100 steps) --- ## ๐Ÿ“Š Evaluation ### Perplexity & Bits-per-Byte (Held-out Test, n=50) | Metric | Base Model (Qwen2.5-0.5B) | **This Model** | Improvement | |--------|---------------------------|----------------|-------------| | Avg Loss | 2.359 | **1.655** | โ†“29.8% | | Perplexity | 10.58 | **5.23** | โ†“50.6% | | Bits/Byte | 0.698 | **0.490** | โ†“29.8% | ### Per-Genre Evaluation (Domain-specific) | Genre | Base PPL | **Our PPL** | Improvement | |-------|----------|-------------|-------------| | **Quran** | 3.36 | **1.69** | โ†“49.6% | | **Hadith** | 2.63 | **1.40** | โ†“47.0% | | **Tafsir** | 4.34 | **2.45** | โ†“43.7% | | **Fiqh (Classical)** | 5.72 | **2.41** | โ†“57.8% | | **Prose Turath** | 5.21 | **3.02** | โ†“42.0% | All differences are **statistically significant** (bootstrap CI 95%, p < 0.05). ### Tashkeel Robustness - **This model:** 0.154 (lower = more robust to harakat variations) - **Base:** 0.249 ### Generation Speed - **This model:** 9.89 tokens/sec - **Base:** 8.37 tokens/sec (+18% faster, due to smaller vocabulary) ### Diversity (Distinct-n) | Model | Distinct-1 | Distinct-2 | Distinct-3 | |-------|------------|------------|------------| | This model | 0.878 | **0.989** | 1.0 | | Base | 0.880 | 0.957 | 1.0 | --- ## ๐Ÿ’ก Generation Examples ### Hadith Continuation > **Input:** ุญูŽุฏูŽู‘ุซูŽู†ูŽุง ุณููู’ูŠูŽุงู†ู ุนูŽู†ู ุงู„ุฒูู‘ู‡ู’ุฑููŠูู‘ > **Output:** ุญูŽุฏูŽู‘ุซูŽู†ูŽุง ุณููู’ูŠูŽุงู†ู ุนูŽู†ู ุงู„ุฒูู‘ู‡ู’ุฑููŠูู‘ ุนูŽู†ู’ ุนูุจูŽูŠู’ุฏู ุงู„ู„ูŽู‘ู‡ู ุจู’ู†ู ุนูŽุจู’ุฏู ุงู„ู„ูŽู‘ู‡ู ุจู’ู†ู ุนูุชู’ุจูŽุฉูŽ ุนูŽู†ู’ ุฃูŽุจููŠู‡ู ู‚ูŽุงู„ูŽ : ูƒูŽุงู†ูŽ ุฑ... ### Fiqh Ruling > **Input:** ูˆูŽุงู„ุฏูŽู‘ู„ููŠู„ู ุนูŽู„ูŽู‰ ูˆูุฌููˆุจู ู‡ูŽุฐูู‡ู ุงู„ู’ู…ูŽุณู’ุฃูŽู„ูŽุฉู ุนูู†ู’ุฏูŽ ุงู„ู’ููู‚ูŽู‡ูŽุงุกู > **Output:** ูˆูŽุงู„ุฏูŽู‘ู„ููŠู„ู ุนูŽู„ูŽู‰ ูˆูุฌููˆุจู ู‡ูŽุฐูู‡ู ุงู„ู’ู…ูŽุณู’ุฃูŽู„ูŽุฉู ุนูู†ู’ุฏูŽ ุงู„ู’ููู‚ูŽู‡ูŽุงุกู ุฃูŽู†ูŽู‘ู‡ูŽุง ุฅูุฐูŽุง ูƒูŽุงู†ูŽุชู’ ู…ูุญูŽุฑูŽู‘ู…ูŽุฉูŒ ููŽุฅูู†ูŽู‘ู‡ูŽุง ุชูŽูƒููˆู†ู ู…ูุญูŽุฑูŽู‘ู…ูŽุฉูŒ ุจูุงู„ู’... ### Tafsir Style > **Input:** ู‚ูŽูˆู’ู„ูู‡ู ุชูŽุนูŽุงู„ูŽู‰ ูˆูŽุงุตู’ุจูุฑู’ ู†ูŽูู’ุณูŽูƒูŽ ู…ูŽุนูŽ ุงู„ูŽู‘ุฐููŠู†ูŽ > **Output:** ู‚ูŽูˆู’ู„ูู‡ู ุชูŽุนูŽุงู„ูŽู‰ ูˆูŽุงุตู’ุจูุฑู’ ู†ูŽูู’ุณูŽูƒูŽ ู…ูŽุนูŽ ุงู„ูŽู‘ุฐููŠู†ูŽ ูŠูŽุนู’ู†ููŠ ุจูุฃูŽุฑู’ุญูŽุงู…ูู‡ูู…ู’ ุฃูŽู†ููุณูŽู‡ูู…ู’ ูˆูŽุฃูŽุฒู’ูˆูŽุฌูŽู‡ูู…ู’ ูˆูŽุฐูุฑูู‘ูŠูŽู‘ุชูŽู‡ูู…ู’ ุฅูู†ูŽู‘... --- ## ๐Ÿš€ Usage ### Requirements ```bash pip install transformers>=4.37.0 torch ``` ### Load Model ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_name = "Ik45/qwen2.5-0.5b-arabic-classic-pruned" model = AutoModelForCausalLM.from_pretrained( model_name, torch_dtype="auto", device_map="auto", trust_remote_code=True ) tokenizer = AutoTokenizer.from_pretrained( model_name, trust_remote_code=True ) # Generate prompt = "ู‚ูŽุงู„ูŽ ุงู„ุฅู…ุงู… ู…ุงู„ูƒ ุฑุญู…ู‡ ุงู„ู„ู‡:" inputs = tokenizer(prompt, return_tensors="pt").to(model.device) outputs = model.generate( **inputs, max_new_tokens=128, do_sample=True, temperature=0.7, top_p=0.9 ) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` --- ## โš ๏ธ Important Notes 1. **Base Model** โ€” This is a *base* model (not instruct/chat). Not recommended for direct dialogue. Use SFT/DPO first if you want to use it as a chatbot. 2. **Domain Specific** โ€” The model is highly specialized for classical Arabic text. Performance on Modern Standard Arabic or other languages may degrade due to aggressive vocabulary pruning. 3. **Context Length** โ€” Trained on 1,024 tokens, but the architecture supports up to 32,768 tokens. For extremely long contexts, additional fine-tuning with RoPE scaling is recommended. 4. **Vocabulary Pruning** โ€” Since 80% of the original Qwen vocabulary was removed, tokenization of non-Arabic text will produce longer sequences (fallback to subword/character-level). --- ## ๐Ÿ“š References ### Papers & Techniques - **Leaf-Based Vocabulary Pruning:** Purason et al., *"Teaching Old Tokenizers New Words"*, arXiv:2512.03989v2 (2026). - **GaLore:** Zhao et al., *"GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection"*, arXiv:2403.03507. - **Qwen2.5:** Qwen Team, *"Qwen2.5: A Party of Foundation Models"*, 2024. ### Dataset - [Shamela/Waqfeya](https://waqfeya.com/) โ€” Digital collection of classical Islamic books. - Pretraining dataset: ~300 million tokens from various genres (Quran, Hadith, Tafsir, Fiqh, Aqidah, Nahwu, Balagha). --- ## ๐Ÿ™ Acknowledgments - [Qwen Team](https://huggingface.co/Qwen) for the Qwen2.5-0.5B base model. - [Purason et al.](https://arxiv.org/abs/2512.03989) for the leaf-based vocabulary pruning method. - [Shamela/Waqfeya](https://waqfeya.com/) for the classical Arabic book dataset. --- ## ๐Ÿ“„ License This model is released under the **Apache 2.0** license, same as the base Qwen2.5-0.5B model. """