jlind456 commited on
Commit
24e8996
·
verified ·
1 Parent(s): 1d1baca

Add model card YAML metadata to README.md

Browse files
Files changed (1) hide show
  1. README.md +28 -115
README.md CHANGED
@@ -1,129 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # Jason AI Twin
2
 
3
  Local, offline digital twin model and voice cloning repository.
4
 
5
  ## Models and Voice Files
6
 
7
- * **digital_twin_q4.gguf:** The quantized local LLM brain.
8
- * **amy.onnx / amy.onnx.json:** Cloned voice model for Piper TTS.
9
- * **my_voice_clean.wav:** Clean source audio reference used to clone the speaker voice with Coqui.
 
 
10
 
11
  ---
12
 
13
- # llama.cpp/example/tts
14
- This example demonstrates the Text To Speech feature. It uses a
15
- [model](https://www.outeai.com/blog/outetts-0.2-500m) from
16
- [outeai](https://www.outeai.com/).
17
-
18
  ## Quickstart
19
- If you have built llama.cpp with SSL support you can simply run the
20
- following command and the required models will be downloaded automatically:
21
- ```console
22
- $ build/bin/llama-tts --tts-oute-default -p "Hello world" && aplay output.wav
23
- ```
24
- For details about the models and how to convert them to the required format
25
- see the following sections.
26
 
27
- ### Model conversion
28
- Checkout or download the model that contains the LLM model:
29
- ```console
30
- $ pushd models
31
- $ git clone --branch main --single-branch --depth 1 https://huggingface.co/OuteAI/OuteTTS-0.2-500M
32
- $ cd OuteTTS-0.2-500M && git lfs install && git lfs pull
33
- $ popd
34
- ```
35
- Convert the model to .gguf format:
36
- ```console
37
- (venv) python convert_hf_to_gguf.py models/OuteTTS-0.2-500M \
38
- --outfile models/outetts-0.2-0.5B-f16.gguf --outtype f16
39
  ```
40
- The generated model will be `models/outetts-0.2-0.5B-f16.gguf`.
41
-
42
- We can optionally quantize this to Q8_0 using the following command:
43
- ```console
44
- $ build/bin/llama-quantize models/outetts-0.2-0.5B-f16.gguf \
45
- models/outetts-0.2-0.5B-q8_0.gguf q8_0
46
- ```
47
- The quantized model will be `models/outetts-0.2-0.5B-q8_0.gguf`.
48
-
49
- Next we do something similar for the audio decoder. First download or checkout
50
- the model for the voice decoder:
51
- ```console
52
- $ pushd models
53
- $ git clone --branch main --single-branch --depth 1 https://huggingface.co/novateur/WavTokenizer-large-speech-75token
54
- $ cd WavTokenizer-large-speech-75token && git lfs install && git lfs pull
55
- $ popd
56
- ```
57
- This model file is a PyTorch checkpoint (.ckpt) and we first need to convert it to
58
- huggingface format:
59
- ```console
60
- (venv) python tools/tts/convert_pt_to_hf.py \
61
- models/WavTokenizer-large-speech-75token/wavtokenizer_large_speech_320_24k.ckpt
62
- ...
63
- Model has been successfully converted and saved to models/WavTokenizer-large-speech-75token/model.safetensors
64
- Metadata has been saved to models/WavTokenizer-large-speech-75token/index.json
65
- Config has been saved to models/WavTokenizer-large-speech-75tokenconfig.json
66
- ```
67
- Then we can convert the huggingface format to gguf:
68
- ```console
69
- (venv) python convert_hf_to_gguf.py models/WavTokenizer-large-speech-75token \
70
- --outfile models/wavtokenizer-large-75-f16.gguf --outtype f16
71
- ...
72
- INFO:hf-to-gguf:Model successfully exported to models/wavtokenizer-large-75-f16.gguf
73
- ```
74
-
75
- ### Running the example
76
 
77
- With both of the models generated, the LLM model and the voice decoder model,
78
- we can run the example:
79
- ```console
80
- $ build/bin/llama-tts -m ./models/outetts-0.2-0.5B-q8_0.gguf \
81
- -mv ./models/wavtokenizer-large-75-f16.gguf \
82
- -p "Hello world"
83
- ...
84
- main: audio written to file 'output.wav'
85
- ```
86
- The output.wav file will contain the audio of the prompt. This can be heard
87
- by playing the file with a media player. On Linux the following command will
88
- play the audio:
89
- ```console
90
- $ aplay output.wav
91
- ```
92
-
93
- ### Running the example with llama-server
94
- Running this example with `llama-server` is also possible and requires two
95
- server instances to be started. One will serve the LLM model and the other
96
- will serve the voice decoder model.
97
-
98
- The LLM model server can be started with the following command:
99
- ```console
100
- $ ./build/bin/llama-server -m ./models/outetts-0.2-0.5B-q8_0.gguf --port 8020
101
- ```
102
-
103
- And the voice decoder model server can be started using:
104
- ```console
105
- ./build/bin/llama-server -m ./models/wavtokenizer-large-75-f16.gguf --port 8021 --embeddings --pooling none
106
- ```
107
-
108
- Then we can run [tts-outetts.py](tts-outetts.py) to generate the audio.
109
-
110
- First create a virtual environment for python and install the required
111
- dependencies (this in only required to be done once):
112
- ```console
113
- $ python3 -m venv venv
114
- $ source venv/bin/activate
115
- (venv) pip install requests numpy
116
- ```
117
-
118
- And then run the python script using:
119
- ```console
120
- (venv) python ./tools/tts/tts-outetts.py http://localhost:8020 http://localhost:8021 "Hello world"
121
- spectrogram generated: n_codes: 90, n_embd: 1282
122
- converting to audio ...
123
- audio generated: 28800 samples
124
- audio written to file "output.wav"
125
- ```
126
- And to play the audio we can again use aplay or any other media player:
127
- ```console
128
- $ aplay output.wav
129
  ```
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: mit
5
+ library_name: gguf
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - gguf
9
+ - conversational
10
+ - text-generation
11
+ - text-to-speech
12
+ - voice-cloning
13
+ - offline
14
+ - digital-twin
15
+ ---
16
+
17
  # Jason AI Twin
18
 
19
  Local, offline digital twin model and voice cloning repository.
20
 
21
  ## Models and Voice Files
22
 
23
+ * **`digital_twin_q4.gguf`**: The quantized local LLM brain (`jason_twin`).
24
+ * **`cloned_output.wav`**: Cloned voice reference audio for XTTS v2 text-to-speech.
25
+ * **`my_voice_clean.wav`**: Clean source audio reference for voice synthesis.
26
+ * **`amy.onnx` / `amy.onnx.json`**: Cloned voice model for Piper TTS.
27
+ * **`chat_twin.py`**: Interactive CLI interface for AI Twin chat and speech.
28
 
29
  ---
30
 
 
 
 
 
 
31
  ## Quickstart
 
 
 
 
 
 
 
32
 
33
+ ### Running the AI Twin CLI:
34
+ ```bash
35
+ python3 chat_twin.py
 
 
 
 
 
 
 
 
 
36
  ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
 
38
+ ### Running with Ollama:
39
+ ```bash
40
+ ollama create jason_twin -f Modelfile
41
+ ollama run jason_twin
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42
  ```