nemotron-speech-streaming-en-0.6b β ONNX, calibrated static int8
A calibration-based static-int8 quantization of the ONNX export of nvidia/nemotron-speech-streaming-en-0.6b (English streaming speech recognition, 0.6B parameters).
- Peak inference memory β 0.8 GB (the fp32 ONNX export runs β 2.6 GB).
- Word-error rate within noise of fp32 on standard read-speech material; whispered/very-quiet speech degrades somewhat relative to fp32.
- Files:
encoder.onnx+encoder.onnx.data,decoder_joint.onnx,tokenizer.model. - Runs with ONNX Runtime; the encoder is cache-aware for incremental/streaming inference.
Licensed under the NVIDIA Open Model License β see LICENSE and NOTICE.
Base model Β© NVIDIA Corporation.
Model tree for potgieterdl/nemotron-speech-streaming-en-0.6b-onnx-int8
Base model
nvidia/nemotron-speech-streaming-en-0.6b