nemotron-speech-streaming-en-0.6b β€” ONNX, calibrated static int8

A calibration-based static-int8 quantization of the ONNX export of nvidia/nemotron-speech-streaming-en-0.6b (English streaming speech recognition, 0.6B parameters).

  • Peak inference memory β‰ˆ 0.8 GB (the fp32 ONNX export runs β‰ˆ 2.6 GB).
  • Word-error rate within noise of fp32 on standard read-speech material; whispered/very-quiet speech degrades somewhat relative to fp32.
  • Files: encoder.onnx + encoder.onnx.data, decoder_joint.onnx, tokenizer.model.
  • Runs with ONNX Runtime; the encoder is cache-aware for incremental/streaming inference.

Licensed under the NVIDIA Open Model License β€” see LICENSE and NOTICE. Base model Β© NVIDIA Corporation.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for potgieterdl/nemotron-speech-streaming-en-0.6b-onnx-int8

Quantized
(22)
this model