Oído int4: 8.3 MB speech recognition for the ESP32-S3

Animated demo: the transcripts are Oído's output (sped up). Footage from a physical board is coming.

¡Oído! is Spanish kitchen slang for heard, got it.

This is the compact profile of Oído: open-vocabulary English speech recognition that runs entirely on an ESP32-S3, with no cloud and no NPU. At 8.3 MB it leaves a 6 MB app partition free on a 16 MB flash module for your own application code (esp32/firmware/partitions_nemo4.csv).

LibriSpeech WER (%) test-clean test-other Size
This model (int4, greedy, on-chip arithmetic) 4.61 9.98 8.3 MB
Oído int8 3.70 8.23 14.0 MB
Espressif MultiNet7 on the same chip (ESP-SR benchmark) 8.5 21.3 2.9 MB

The int4 weights also run about 10% faster than int8 (estimated RTF 0.68–0.82 from exact QEMU instruction counts; not yet measured on silicon).

Need lower latency? The streaming variant shows text while you speak and delivers the final text about 1.1–1.4 s after you stop (estimated), at some cost in accuracy.

Which Oído model?

One model per language; the models that stream also run in full-context (utterance) mode.

Model Language Modes Size License
oido-ctc-small-int8 English utterance only (best accuracy) 14.0 MB (+1.3 MB LM) CC-BY-4.0
oido-ctc-small-int4 English utterance only (smallest) 8.3 MB CC-BY-SA-4.0
oido-ctc-small-stream-int8 English streaming + full-context (low latency) 14.0 MB CC-BY-SA-4.0
oido-es-ctc-small-int8 Spanish streaming + full-context in one file 14.0 MB (+1.3 MB LM) CC-BY-4.0

Streaming-capable files stream automatically in the firmware's microphone mode and in live_demo.py.

Use

git clone https://github.com/lokutor-ai/oido && cd oido
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo4.tnm      # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone

How it was made

We took NVIDIA's stt_en_conformer_ctc_small (CC-BY-4.0) and fine-tuned it with 4-bit quantization-aware training for 8,000 steps. Linear layers inside the Conformer blocks use int4 per-channel weights; the front end and output head use int8; activations and attention are int8. The training used public corpora:

  • Common Voice 17 and VoxPopuli (CC0);
  • LibriSpeech, MLS English, AMI and VCTK (CC-BY-4.0);
  • People's Speech, with transcripts regenerated by NVIDIA parakeet-tdt-0.6b-v2;
  • OpenSLR 70 and 83 (CC-BY-SA-4.0).

License

This model is released under CC-BY-SA-4.0: it is derived from NVIDIA's CC-BY-4.0 model and trained on data that includes share-alike sources. The Oído engine and firmware are GPLv3, with commercial licenses available from Lokutor.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lokutor-ai/oido-ctc-small-int4

Finetuned
(5)
this model

Datasets used to train lokutor-ai/oido-ctc-small-int4

Evaluation results