Automatic Speech Recognition
GGUF
NeMo
parakeet.cpp
asr
parakeet
ggml
cpp-inference
speaker-diarization
Instructions to use mudler/parakeet-cpp-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mudler/parakeet-cpp-gguf with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mudler/parakeet-cpp-gguf") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
README: describe the bundles, their licences and credits
Browse files
README.md
CHANGED
|
@@ -29,6 +29,8 @@ base_model:
|
|
| 29 |
- moondream/parakeet-ultra
|
| 30 |
- moondream/parakeet-redux
|
| 31 |
- snakers4/silero-vad
|
|
|
|
|
|
|
| 32 |
---
|
| 33 |
|
| 34 |
# Parakeet GGUF — models for parakeet.cpp
|
|
@@ -250,6 +252,56 @@ Source: the voice-activity head, the mel front end and the subsampler of [moondr
|
|
| 250 |
- They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use `parakeet-cli vad --model redux-vad.gguf --input audio.wav`.
|
| 251 |
- For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.
|
| 252 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 253 |
## Quantization notes
|
| 254 |
|
| 255 |
Quantization is applied **only** to the large linear weights fed directly into `ggml_mul_mat` (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.
|
|
@@ -279,5 +331,6 @@ Licences differ by model, so the front matter says `license: other`. Each file f
|
|
| 279 |
- `realtime_eou_120m-v1-*`: derived from [nvidia/parakeet_realtime_eou_120m-v1](https://huggingface.co/nvidia/parakeet_realtime_eou_120m-v1), governed by the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
|
| 280 |
- `nemotron-3.5-asr-streaming-0.6b-*` and `nemotron-3-diarization-*`: derived from NVIDIA Nemotron models, governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/).
|
| 281 |
- `ultra-*` and `redux-*`: see below.
|
|
|
|
| 282 |
|
| 283 |
The notes below add detail. `ultra-*.gguf` and `redux-*.gguf` are converted from [moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux) by Moondream, which are derived from NVIDIA's [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q8_0 files are dequantized from the ternary weights. `redux-vad.gguf` and `ultra-vad-q8_0.gguf` hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. `silero-vad-*.gguf` is converted from [Silero VAD](https://github.com/snakers4/silero-vad) v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. `nemotron-3-diarization-*.gguf` is derived from nvidia/Nemotron-3-Diarization and `nemotron-3.5-asr-streaming-0.6b-*.gguf` from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/). The parakeet.cpp runtime is MIT-licensed.
|
|
|
|
| 29 |
- moondream/parakeet-ultra
|
| 30 |
- moondream/parakeet-redux
|
| 31 |
- snakers4/silero-vad
|
| 32 |
+
- Wespeaker/wespeaker-voxceleb-resnet34-LM
|
| 33 |
+
- mispeech/ced-small
|
| 34 |
---
|
| 35 |
|
| 36 |
# Parakeet GGUF — models for parakeet.cpp
|
|
|
|
| 252 |
- They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use `parakeet-cli vad --model redux-vad.gguf --input audio.wav`.
|
| 253 |
- For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.
|
| 254 |
|
| 255 |
+
## Bundles
|
| 256 |
+
|
| 257 |
+
A bundle is one GGUF file that holds several of the models above, so you pass one file instead of several. A bundle can hold an ASR model, a Silero VAD, speaker diarization, sound-event tagging (CED) and speaker identification. A bundle is only packaging: each model is copied byte for byte with the type it was published with, and nothing was re-quantized, trained or fine-tuned. These files are converted here, not trained. The format is described in [bundle.md](https://github.com/mudler/parakeet.cpp/blob/master/docs/bundle.md).
|
| 258 |
+
|
| 259 |
+
| File | Size | Contents | Runs on |
|
| 260 |
+
|---|---:|---|---|
|
| 261 |
+
| `parakeet-bundle-small.gguf` | 337.9 MB | tdt_ctc-110m Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 | any backend |
|
| 262 |
+
| `parakeet-bundle-standard.gguf` | 1100.8 MB | tdt-0.6b-v3 Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 | any backend |
|
| 263 |
+
| `parakeet-bundle-moondream-redux.gguf` | 214.6 MB | redux-packed (with its own VAD head), Silero VAD F16 | CPU only, offline only |
|
| 264 |
+
|
| 265 |
+
> A bundle needs a parakeet.cpp build that includes the bundle code (pull request 85). That code is on the `master` branch but not in a release yet (the latest release, v0.5.0, does not have it), so build from source for now. Older builds refuse a bundle with a load error. They never read the wrong weights. The bundles were checked on CPU only: the output of each component equals the output of its single-model file (the transcript, the diarization segments, the CED class scores, the speaker embeddings and the Silero probabilities). GPU backends, streaming ASR from a bundle, macOS and Windows were not tested.
|
| 266 |
+
|
| 267 |
+
```bash
|
| 268 |
+
huggingface-cli download mudler/parakeet-cpp-gguf parakeet-bundle-small.gguf --local-dir models/
|
| 269 |
+
|
| 270 |
+
# List the components, licences and credits (reads only the header)
|
| 271 |
+
build/examples/cli/parakeet-cli info models/parakeet-bundle-small.gguf
|
| 272 |
+
|
| 273 |
+
# Transcribe with the ASR component; add --vad to cut long audio with the Silero component
|
| 274 |
+
build/examples/cli/parakeet-cli transcribe --model models/parakeet-bundle-small.gguf --input audio.wav
|
| 275 |
+
|
| 276 |
+
# Who spoke when, with the diarization component
|
| 277 |
+
build/examples/cli/diarize models/parakeet-bundle-small.gguf meeting.wav
|
| 278 |
+
|
| 279 |
+
# Speaker-attributed transcript with sound events: one file passed for every role
|
| 280 |
+
build/examples/cli/parakeet-cli scene --model models/parakeet-bundle-small.gguf \
|
| 281 |
+
--diar models/parakeet-bundle-small.gguf --sound models/parakeet-bundle-small.gguf \
|
| 282 |
+
--input meeting.wav
|
| 283 |
+
# Add --speakers models/parakeet-bundle-small.gguf --registry people.bin to name known voices
|
| 284 |
+
```
|
| 285 |
+
|
| 286 |
+
The moondream-redux bundle has the ASR and Silero components only. Use `parakeet-cli transcribe --model models/parakeet-bundle-moondream-redux.gguf --input audio.wav --vad`: it cuts at pauses with the Silero component, and `--vad-component asr` uses the Redux head instead. When a bundle has more than one component of a kind, name one with `--component`, `--asr-component`, `--diar-component`, `--sound-component` or `--speakers-component`.
|
| 287 |
+
|
| 288 |
+
**Licences.** A bundle has no single licence, so the header says `other` and each component keeps the licence of the model it was converted from. The credit, the licence link and the changes are in the file header (`parakeet-cli info` shows them), in `NOTICE-parakeet-bundle-<name>.txt` next to each bundle, and the full licence texts are in the [`licenses/`](https://huggingface.co/mudler/parakeet-cpp-gguf/tree/main/licenses) folder of this repo. Keep these notices when you redistribute a bundle.
|
| 289 |
+
|
| 290 |
+
| Component | Model | Licence | Credit |
|
| 291 |
+
|---|---|---|---|
|
| 292 |
+
| ASR (small) | [nvidia/parakeet-tdt_ctc-110m](https://huggingface.co/nvidia/parakeet-tdt_ctc-110m) | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) | NVIDIA |
|
| 293 |
+
| ASR (standard) | [nvidia/parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) | NVIDIA |
|
| 294 |
+
| ASR (moondream-redux) | [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux), derived from parakeet-tdt-0.6b-v3 | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) | Moondream and NVIDIA |
|
| 295 |
+
| Diarization | [nvidia/Nemotron-3-Diarization](https://huggingface.co/nvidia/Nemotron-3-Diarization) | [OpenMDW-1.1](https://openmdw.ai/license/1-1/) | NVIDIA |
|
| 296 |
+
| Sound events | [mispeech/ced-small](https://huggingface.co/mispeech/ced-small) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) (see the note below) | Heinrich Dinkel et al., Xiaomi (mispeech) |
|
| 297 |
+
| Speaker identification | [Wespeaker/wespeaker-voxceleb-resnet34-LM](https://huggingface.co/Wespeaker/wespeaker-voxceleb-resnet34-LM) | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/) (see the note below) | the [WeSpeaker project](https://github.com/wenet-e2e/wespeaker) |
|
| 298 |
+
| VAD | [snakers4/silero-vad](https://github.com/snakers4/silero-vad) | [MIT](https://github.com/snakers4/silero-vad/blob/master/LICENSE) | Copyright (c) 2020-present Silero Team |
|
| 299 |
+
|
| 300 |
+
- **CED:** the `mispeech/ced-*` model cards say Apache-2.0, and the bundle follows them. The upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream. It has not been confirmed with the authors. The CC-BY credit to the authors is kept in the meantime.
|
| 301 |
+
- **WeSpeaker:** the file is converted from `voxceleb_resnet34_LM.onnx` of `Wespeaker/wespeaker-voxceleb-resnet34-LM`, whose card says CC-BY-4.0. The card of the plain `wespeaker-voxceleb-resnet34` says Apache-2.0, but the WeSpeaker project states in its documentation that its VoxCeleb-trained models follow CC-BY-4.0, so the bundle uses CC-BY-4.0 and credits the WeSpeaker project. The speaker models are trained on VoxCeleb. Whether a trained model is derived from its training data is a legal question that this project does not settle.
|
| 302 |
+
- **Changes:** the models are converted to GGUF here and, for the ASR and diarization models, quantized to Q8_0 (the redux-packed ASR component is the published packed file). Nothing was trained or fine-tuned.
|
| 303 |
+
- The end-of-utterance model and the audeering voice-analysis heads are never put in a bundle: their licences do not allow it.
|
| 304 |
+
|
| 305 |
## Quantization notes
|
| 306 |
|
| 307 |
Quantization is applied **only** to the large linear weights fed directly into `ggml_mul_mat` (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.
|
|
|
|
| 331 |
- `realtime_eou_120m-v1-*`: derived from [nvidia/parakeet_realtime_eou_120m-v1](https://huggingface.co/nvidia/parakeet_realtime_eou_120m-v1), governed by the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
|
| 332 |
- `nemotron-3.5-asr-streaming-0.6b-*` and `nemotron-3-diarization-*`: derived from NVIDIA Nemotron models, governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/).
|
| 333 |
- `ultra-*` and `redux-*`: see below.
|
| 334 |
+
- `parakeet-bundle-*`: one licence per component, see [Bundles](#bundles).
|
| 335 |
|
| 336 |
The notes below add detail. `ultra-*.gguf` and `redux-*.gguf` are converted from [moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux) by Moondream, which are derived from NVIDIA's [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q8_0 files are dequantized from the ternary weights. `redux-vad.gguf` and `ultra-vad-q8_0.gguf` hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. `silero-vad-*.gguf` is converted from [Silero VAD](https://github.com/snakers4/silero-vad) v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. `nemotron-3-diarization-*.gguf` is derived from nvidia/Nemotron-3-Diarization and `nemotron-3.5-asr-streaming-0.6b-*.gguf` from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the [OpenMDW License Agreement, version 1.1](https://openmdw.ai/license/1-1/). The parakeet.cpp runtime is MIT-licensed.
|