Parakeet TDT 0.6B v3 路 ONNX (int8)

ONNX build of NVIDIA Parakeet TDT 0.6B v3, a multilingual automatic speech recognition model, quantized to int8 for fast, low latency inference on CPU and on device. Hosted by Palatine, where it powers speech to text in Palatine Notes. Provided as is.

Languages

Multilingual across 25 European languages: English, Spanish, French, German, Bulgarian, Croatian, Czech, Danish, Dutch, Estonian, Finnish, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Swedish, Russian, and Ukrainian.

Files

This repository ships both the int8 quantized and the full precision graphs of the encoder (encoder-model.int8.onnx, and encoder-model.onnx with its encoder-model.onnx.data weights) and the decoder / joint network (decoder_joint-model.int8.onnx, decoder_joint-model.onnx), along with the feature extractor (nemo128.onnx) and the vocabulary (vocab.txt).

License & attribution

Base model 漏 NVIDIA, released under CC-BY-4.0; this ONNX build is redistributed under the same license.

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for PalatineVision/parakeet-tdt-0.6b-v3-onnx

Quantized
(64)
this model