Text-to-Speech
PEFT
Safetensors
tts
speech-synthesis
orpheus
snac
lora
unsloth
yoruba
hausa
igbo
nigerian-pidgin
nigeria
african-languages
multilingual
low-resource
Instructions to use Shinzmann/sorotts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Shinzmann/sorotts with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("/root/.cache/huggingface/hub/models--hypaai--hypaai_orpheus_v5/snapshots/a8786380f8f8c9b1215bc5b299ab740b3df1781d") model = PeftModel.from_pretrained(base_model, "Shinzmann/sorotts") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Studio
How to use Shinzmann/sorotts with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Shinzmann/sorotts to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Shinzmann/sorotts to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Shinzmann/sorotts to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="Shinzmann/sorotts", max_seq_length=2048, )
Training details from the real B200 run: 97.3M trainable (2.86%), 7894 steps, ~31 min
Browse files
README.md
CHANGED
|
@@ -159,8 +159,8 @@ Because NaijaVoices (non-commercial) is included, the model is stamped **`cc-by-
|
|
| 159 |
|
| 160 |
## How it was trained
|
| 161 |
|
| 162 |
-
- **LoRA** r=64, α=64, dropout 0, on all attention and MLP projections (the Orpheus-TTS standard), via **[Unsloth](https://github.com/unslothai/unsloth)**.
|
| 163 |
-
- **2 epochs
|
| 164 |
- Audio is encoded to Orpheus's **7-tokens-per-frame SNAC stream**; each training sequence is
|
| 165 |
`[SOH] voice: text [EOT][EOH] [SOAI][SOS] <snac codes> [EOS][EOAI]`.
|
| 166 |
- **Fully reproducible on Modal**: streaming and SNAC-encoding the data, the LoRA training, and the Hub push all run as one serverless job. See [`modal/finetune_orpheus.py`](https://github.com/Mystique1337/naija-solar/blob/main/modal/finetune_orpheus.py) (train), [`serving_tts.py`](https://github.com/Mystique1337/naija-solar/blob/main/modal/serving_tts.py) (serve), and [`test_sorotts.py`](https://github.com/Mystique1337/naija-solar/blob/main/modal/test_sorotts.py) (samples).
|
|
|
|
| 159 |
|
| 160 |
## How it was trained
|
| 161 |
|
| 162 |
+
- **LoRA** r=64, α=64, dropout 0, on all attention and MLP projections (the Orpheus-TTS standard), via **[Unsloth](https://github.com/unslothai/unsloth)**. Only **97.3M parameters are trained, 2.86%** of the 3.40B-parameter model.
|
| 163 |
+
- **2 epochs = 7,894 steps**, bf16, AdamW-8bit, lr 2e-4, 3% warmup, **total batch 8** (8 per device, 1 gradient-accumulation step), on a single **NVIDIA B200**. The run takes about **31 minutes** (1,894 s, roughly 33 samples per second) and converges to a train loss near 3.5.
|
| 164 |
- Audio is encoded to Orpheus's **7-tokens-per-frame SNAC stream**; each training sequence is
|
| 165 |
`[SOH] voice: text [EOT][EOH] [SOAI][SOS] <snac codes> [EOS][EOAI]`.
|
| 166 |
- **Fully reproducible on Modal**: streaming and SNAC-encoding the data, the LoRA training, and the Hub push all run as one serverless job. See [`modal/finetune_orpheus.py`](https://github.com/Mystique1337/naija-solar/blob/main/modal/finetune_orpheus.py) (train), [`serving_tts.py`](https://github.com/Mystique1337/naija-solar/blob/main/modal/serving_tts.py) (serve), and [`test_sorotts.py`](https://github.com/Mystique1337/naija-solar/blob/main/modal/test_sorotts.py) (samples).
|