--- license: apache-2.0 base_model: unsloth/Qwen3.5-4B-Base tags: [tajik, qwen3.5, continual-pretraining, cpt, merged] language: [tg] --- # Qwen3.5-4B Tajik — CPT epoch-1, merged Continual-pretraining (CPT) adaptation of **unsloth/Qwen3.5-4B-Base** to **Tajik** (тоҷикӣ), epoch 1, with the LoRA adapter merged into the base weights (BF16). Equivalent to loading `Tohirju/qwen35-4b-tajik-cpt-e1` on the base. - **What it is:** the base Qwen3.5-4B after 1 epoch of Tajik continual pre-training (~314M Tajik tokens). - **Use:** a stronger Tajik base for generation and as the starting point for the next CPT epoch or for SFT. - **Tajik MCQ (Zehnlab, full ~18,970 Q, loglikelihood):** ~48.6% avg (base 44.7%). See the eval repo. - **Note:** this is a base/CPT model — it *completes* text, it does not follow chat instructions (that is the SFT stage). - Merge verified by logit parity against the adapter (max|Δlogit| < 1e0). Sibling adapters: `Tohirju/qwen35-4b-tajik-cpt-e1`, `-e2`, `-e3`. Gated: download requires manual approval. Released by Saidzoda Lab / Saidzoda Engineering Company.