audiosronnx-gtcrn

ONNX weights for the gtcrn denoise engine in audiosronnx — a mirror of the streaming GTCRN graph from Xiaobin-Rong/gtcrn (MIT).

GTCRN (Grouped Temporal Convolutional Recurrent Network) is an ultra-light 16 kHz speech denoiser — 23.7 K parameters, 33 MMACs/s, ~0.5 MB — intended for embedded and on-device use. The graph is stateful and consumes one STFT frame at a time:

mix[1, 257, 1, 2], conv_cache, tra_cache, inter_cache
    -> enh[1, 257, 1, 2], conv_cache_out, tra_cache_out, inter_cache_out

The ERB filterbank and subband feature extraction live inside the graph; only a sqrt-Hann STFT/ISTFT runs outside it, in numpy, so inference is onnxruntime-only.

File Rate Size
gtcrn_simple.onnx 16 kHz 0.54 MB

Usage

from audiosronnx import load_denoise

dn = load_denoise("gtcrn")
clean, rate = dn.denoise("noisy.wav")   # -> (float32 mono, 16000)

License

MIT, inherited from the upstream GTCRN release.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support