mudler commited on
Commit
741158a
·
verified ·
1 Parent(s): 61a970c

README: describe the bundles, their licences and credits

Browse files
Files changed (1) hide show
  1. README.md +53 -0
README.md CHANGED
@@ -29,6 +29,8 @@ base_model:
29
  - moondream/parakeet-ultra
30
  - moondream/parakeet-redux
31
  - snakers4/silero-vad
 
 
32
  ---
33
 
34
  # Parakeet GGUF — models for parakeet.cpp
@@ -250,6 +252,56 @@ Source: the voice-activity head, the mel front end and the subsampler of [moondr
250
  - They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use `parakeet-cli vad --model redux-vad.gguf --input audio.wav`.
251
  - For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.
252
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
253
  ## Quantization notes
254
 
255
  Quantization is applied **only** to the large linear weights fed directly into `ggml_mul_mat` (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.
@@ -279,5 +331,6 @@ Licences differ by model, so the front matter says `license: other`. Each file f
279
  - `realtime_eou_120m-v1-*`: derived from [nvidia/parakeet_realtime_eou_120m-v1](https://huggingface.co/nvidia/parakeet_realtime_eou_120m-v1), governed by the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
280
  - `nemotron-3.5-asr-streaming-0.6b-*` and `nemotron-3-diarization-*`: derived from NVIDIA Nemotron models, governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/).
281
  - `ultra-*` and `redux-*`: see below.
 
282
 
283
  The notes below add detail. `ultra-*.gguf` and `redux-*.gguf` are converted from [moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux) by Moondream, which are derived from NVIDIA's [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q8_0 files are dequantized from the ternary weights. `redux-vad.gguf` and `ultra-vad-q8_0.gguf` hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. `silero-vad-*.gguf` is converted from [Silero VAD](https://github.com/snakers4/silero-vad) v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. `nemotron-3-diarization-*.gguf` is derived from nvidia/Nemotron-3-Diarization and `nemotron-3.5-asr-streaming-0.6b-*.gguf` from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/). The parakeet.cpp runtime is MIT-licensed.
 
29
  - moondream/parakeet-ultra
30
  - moondream/parakeet-redux
31
  - snakers4/silero-vad
32
+ - Wespeaker/wespeaker-voxceleb-resnet34-LM
33
+ - mispeech/ced-small
34
  ---
35
 
36
  # Parakeet GGUF — models for parakeet.cpp
 
252
  - They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use `parakeet-cli vad --model redux-vad.gguf --input audio.wav`.
253
  - For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.
254
 
255
+ ## Bundles
256
+
257
+ A bundle is one GGUF file that holds several of the models above, so you pass one file instead of several. A bundle can hold an ASR model, a Silero VAD, speaker diarization, sound-event tagging (CED) and speaker identification. A bundle is only packaging: each model is copied byte for byte with the type it was published with, and nothing was re-quantized, trained or fine-tuned. These files are converted here, not trained. The format is described in [bundle.md](https://github.com/mudler/parakeet.cpp/blob/master/docs/bundle.md).
258
+
259
+ | File | Size | Contents | Runs on |
260
+ |---|---:|---|---|
261
+ | `parakeet-bundle-small.gguf` | 337.9 MB | tdt_ctc-110m Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 | any backend |
262
+ | `parakeet-bundle-standard.gguf` | 1100.8 MB | tdt-0.6b-v3 Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 | any backend |
263
+ | `parakeet-bundle-moondream-redux.gguf` | 214.6 MB | redux-packed (with its own VAD head), Silero VAD F16 | CPU only, offline only |
264
+
265
+ > A bundle needs a parakeet.cpp build that includes the bundle code (pull request 85). That code is on the `master` branch but not in a release yet (the latest release, v0.5.0, does not have it), so build from source for now. Older builds refuse a bundle with a load error. They never read the wrong weights. The bundles were checked on CPU only: the output of each component equals the output of its single-model file (the transcript, the diarization segments, the CED class scores, the speaker embeddings and the Silero probabilities). GPU backends, streaming ASR from a bundle, macOS and Windows were not tested.
266
+
267
+ ```bash
268
+ huggingface-cli download mudler/parakeet-cpp-gguf parakeet-bundle-small.gguf --local-dir models/
269
+
270
+ # List the components, licences and credits (reads only the header)
271
+ build/examples/cli/parakeet-cli info models/parakeet-bundle-small.gguf
272
+
273
+ # Transcribe with the ASR component; add --vad to cut long audio with the Silero component
274
+ build/examples/cli/parakeet-cli transcribe --model models/parakeet-bundle-small.gguf --input audio.wav
275
+
276
+ # Who spoke when, with the diarization component
277
+ build/examples/cli/diarize models/parakeet-bundle-small.gguf meeting.wav
278
+
279
+ # Speaker-attributed transcript with sound events: one file passed for every role
280
+ build/examples/cli/parakeet-cli scene --model models/parakeet-bundle-small.gguf \
281
+ --diar models/parakeet-bundle-small.gguf --sound models/parakeet-bundle-small.gguf \
282
+ --input meeting.wav
283
+ # Add --speakers models/parakeet-bundle-small.gguf --registry people.bin to name known voices
284
+ ```
285
+
286
+ The moondream-redux bundle has the ASR and Silero components only. Use `parakeet-cli transcribe --model models/parakeet-bundle-moondream-redux.gguf --input audio.wav --vad`: it cuts at pauses with the Silero component, and `--vad-component asr` uses the Redux head instead. When a bundle has more than one component of a kind, name one with `--component`, `--asr-component`, `--diar-component`, `--sound-component` or `--speakers-component`.
287
+
288
+ **Licences.** A bundle has no single licence, so the header says `other` and each component keeps the licence of the model it was converted from. The credit, the licence link and the changes are in the file header (`parakeet-cli info` shows them), in `NOTICE-parakeet-bundle-<name>.txt` next to each bundle, and the full licence texts are in the [`licenses/`](https://huggingface.co/mudler/parakeet-cpp-gguf/tree/main/licenses) folder of this repo. Keep these notices when you redistribute a bundle.
289
+
290
+ | Component | Model | Licence | Credit |
291
+ |---|---|---|---|
292
+ | ASR (small) | [nvidia/parakeet-tdt_ctc-110m](https://huggingface.co/nvidia/parakeet-tdt_ctc-110m) | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) | NVIDIA |
293
+ | ASR (standard) | [nvidia/parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) | NVIDIA |
294
+ | ASR (moondream-redux) | [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux), derived from parakeet-tdt-0.6b-v3 | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) | Moondream and NVIDIA |
295
+ | Diarization | [nvidia/Nemotron-3-Diarization](https://huggingface.co/nvidia/Nemotron-3-Diarization) | [OpenMDW-1.1](https://openmdw.ai/license/1-1/) | NVIDIA |
296
+ | Sound events | [mispeech/ced-small](https://huggingface.co/mispeech/ced-small) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) (see the note below) | Heinrich Dinkel et al., Xiaomi (mispeech) |
297
+ | Speaker identification | [Wespeaker/wespeaker-voxceleb-resnet34-LM](https://huggingface.co/Wespeaker/wespeaker-voxceleb-resnet34-LM) | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) (see the note below) | the [WeSpeaker project](https://github.com/wenet-e2e/wespeaker) |
298
+ | VAD | [snakers4/silero-vad](https://github.com/snakers4/silero-vad) | [MIT](https://github.com/snakers4/silero-vad/blob/master/LICENSE) | Copyright (c) 2020-present Silero Team |
299
+
300
+ - **CED:** the `mispeech/ced-*` model cards say Apache-2.0, and the bundle follows them. The upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream. It has not been confirmed with the authors. The CC-BY credit to the authors is kept in the meantime.
301
+ - **WeSpeaker:** the file is converted from `voxceleb_resnet34_LM.onnx` of `Wespeaker/wespeaker-voxceleb-resnet34-LM`, whose card says CC-BY-4.0. The card of the plain `wespeaker-voxceleb-resnet34` says Apache-2.0, but the WeSpeaker project states in its documentation that its VoxCeleb-trained models follow CC-BY-4.0, so the bundle uses CC-BY-4.0 and credits the WeSpeaker project. The speaker models are trained on VoxCeleb. Whether a trained model is derived from its training data is a legal question that this project does not settle.
302
+ - **Changes:** the models are converted to GGUF here and, for the ASR and diarization models, quantized to Q8_0 (the redux-packed ASR component is the published packed file). Nothing was trained or fine-tuned.
303
+ - The end-of-utterance model and the audeering voice-analysis heads are never put in a bundle: their licences do not allow it.
304
+
305
  ## Quantization notes
306
 
307
  Quantization is applied **only** to the large linear weights fed directly into `ggml_mul_mat` (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.
 
331
  - `realtime_eou_120m-v1-*`: derived from [nvidia/parakeet_realtime_eou_120m-v1](https://huggingface.co/nvidia/parakeet_realtime_eou_120m-v1), governed by the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
332
  - `nemotron-3.5-asr-streaming-0.6b-*` and `nemotron-3-diarization-*`: derived from NVIDIA Nemotron models, governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/).
333
  - `ultra-*` and `redux-*`: see below.
334
+ - `parakeet-bundle-*`: one licence per component, see [Bundles](#bundles).
335
 
336
  The notes below add detail. `ultra-*.gguf` and `redux-*.gguf` are converted from [moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux) by Moondream, which are derived from NVIDIA's [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q8_0 files are dequantized from the ternary weights. `redux-vad.gguf` and `ultra-vad-q8_0.gguf` hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. `silero-vad-*.gguf` is converted from [Silero VAD](https://github.com/snakers4/silero-vad) v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. `nemotron-3-diarization-*.gguf` is derived from nvidia/Nemotron-3-Diarization and `nemotron-3.5-asr-streaming-0.6b-*.gguf` from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/). The parakeet.cpp runtime is MIT-licensed.