Instructions to use kyutai/pocket-tts-without-voice-cloning with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use kyutai/pocket-tts-without-voice-cloning with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("kyutai/pocket-tts-without-voice-cloning") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Add pipeline tag and library name to metadata
Browse filesHi! I'm Niels from the Hugging Face community team.
I've opened this PR to improve the model card's discoverability. I've added the `text-to-speech` pipeline tag and the `library_name: pocket-tts` to the metadata. I've also included a link to the project page in the header. These changes help users find the model more easily and understand how to use it with the associated library.
The rest of the content remains unchanged as it already provides excellent documentation.
README.md
CHANGED
|
@@ -1,17 +1,28 @@
|
|
| 1 |
---
|
| 2 |
-
license: cc-by-4.0
|
| 3 |
language:
|
| 4 |
- en
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
extra_gated_fields:
|
| 7 |
Company or university if applicable: text
|
| 8 |
I want to use this model for:
|
| 9 |
type: select
|
| 10 |
-
options:
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
---
|
|
|
|
| 15 |
# Pocket TTS
|
| 16 |
|
| 17 |

|
|
@@ -22,6 +33,7 @@ Forget about the hassle of using GPUs and web APIs serving TTS models. With Kyut
|
|
| 22 |
Supports Python 3.10, 3.11, 3.12, 3.13 and 3.14. Requires PyTorch 2.5+. Does not require the gpu version of PyTorch.
|
| 23 |
|
| 24 |
[🔊 Demo](https://kyutai.org/tts) |
|
|
|
|
| 25 |
[🐱💻GitHub Repository](https://github.com/kyutai-labs/pocket-tts) |
|
| 26 |
[🤗 Hugging Face Model Card](https://huggingface.co/kyutai/pocket-tts) |
|
| 27 |
[📄 Paper](https://arxiv.org/abs/2509.06926) |
|
|
@@ -137,4 +149,4 @@ You can find development instructions in the [CONTRIBUTING.md](https://github.co
|
|
| 137 |
|
| 138 |
## Prohibited use
|
| 139 |
|
| 140 |
-
Use of our model must comply with all applicable laws and regulations and must not result in, involve, or facilitate any illegal, harmful, deceptive, fraudulent, or unauthorized activity. Prohibited uses include, without limitation, voice impersonation or cloning without explicit and lawful consent; misinformation, disinformation, or deception (including fake news, fraudulent calls, or presenting generated content as genuine recordings of real people or events); and the generation of unlawful, harmful, libelous, abusive, harassing, discriminatory, hateful, or privacy-invasive content. We disclaim all liability for any non-compliant use.
|
|
|
|
| 1 |
---
|
|
|
|
| 2 |
language:
|
| 3 |
- en
|
| 4 |
+
license: cc-by-4.0
|
| 5 |
+
pipeline_tag: text-to-speech
|
| 6 |
+
library_name: pocket-tts
|
| 7 |
+
extra_gated_prompt: 'Prohibited use: Use of our model must comply with all applicable
|
| 8 |
+
laws and regulations and must not result in, involve, or facilitate any illegal,
|
| 9 |
+
harmful, deceptive, fraudulent, or unauthorized activity. Prohibited uses include,
|
| 10 |
+
without limitation, voice impersonation or cloning without explicit and lawful consent;
|
| 11 |
+
misinformation, disinformation, or deception (including fake news, fraudulent calls,
|
| 12 |
+
or presenting generated content as genuine recordings of real people or events);
|
| 13 |
+
and the generation of unlawful, harmful, libelous, abusive, harassing, discriminatory,
|
| 14 |
+
hateful, or privacy-invasive content. We disclaim all liability for any non-compliant
|
| 15 |
+
use.'
|
| 16 |
extra_gated_fields:
|
| 17 |
Company or university if applicable: text
|
| 18 |
I want to use this model for:
|
| 19 |
type: select
|
| 20 |
+
options:
|
| 21 |
+
- Work
|
| 22 |
+
- Studies
|
| 23 |
+
- Fun
|
| 24 |
---
|
| 25 |
+
|
| 26 |
# Pocket TTS
|
| 27 |
|
| 28 |

|
|
|
|
| 33 |
Supports Python 3.10, 3.11, 3.12, 3.13 and 3.14. Requires PyTorch 2.5+. Does not require the gpu version of PyTorch.
|
| 34 |
|
| 35 |
[🔊 Demo](https://kyutai.org/tts) |
|
| 36 |
+
[🌐 Project Page](https://iclr-continuous-audio-language-models.github.io) |
|
| 37 |
[🐱💻GitHub Repository](https://github.com/kyutai-labs/pocket-tts) |
|
| 38 |
[🤗 Hugging Face Model Card](https://huggingface.co/kyutai/pocket-tts) |
|
| 39 |
[📄 Paper](https://arxiv.org/abs/2509.06926) |
|
|
|
|
| 149 |
|
| 150 |
## Prohibited use
|
| 151 |
|
| 152 |
+
Use of our model must comply with all applicable laws and regulations and must not result in, involve, or facilitate any illegal, harmful, deceptive, fraudulent, or unauthorized activity. Prohibited uses include, without limitation, voice impersonation or cloning without explicit and lawful consent; misinformation, disinformation, or deception (including fake news, fraudulent calls, or presenting generated content as genuine recordings of real people or events); and the generation of unlawful, harmful, libelous, abusive, harassing, discriminatory, hateful, or privacy-invasive content. We disclaim all liability for any non-compliant use.
|