Oído int4: 8.3 MB speech recognition for the ESP32-S3
Animated demo: the transcripts are Oído's output (sped up). Footage from a physical board is coming.
¡Oído! is Spanish kitchen slang for heard, got it.
This is the compact profile of Oído: open-vocabulary English speech recognition
that runs entirely on an ESP32-S3, with no cloud and no NPU. At 8.3 MB it leaves a 6 MB app partition free on a 16 MB
flash module for your own application code (esp32/firmware/partitions_nemo4.csv).
| LibriSpeech WER (%) | test-clean | test-other | Size |
|---|---|---|---|
| This model (int4, greedy, on-chip arithmetic) | 4.61 | 9.98 | 8.3 MB |
| Oído int8 | 3.70 | 8.23 | 14.0 MB |
| Espressif MultiNet7 on the same chip (ESP-SR benchmark) | 8.5 | 21.3 | 2.9 MB |
The int4 weights also run about 10% faster than int8 (estimated RTF 0.68–0.82 from exact QEMU instruction counts; not yet measured on silicon).
Need lower latency? The streaming variant shows text while you speak and delivers the final text about 1.1–1.4 s after you stop (estimated), at some cost in accuracy.
Which Oído model?
One model per language; the models that stream also run in full-context (utterance) mode.
| Model | Language | Modes | Size | License |
|---|---|---|---|---|
| oido-ctc-small-int8 | English | utterance only (best accuracy) | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 |
| oido-ctc-small-int4 | English | utterance only (smallest) | 8.3 MB | CC-BY-SA-4.0 |
| oido-ctc-small-stream-int8 | English | streaming + full-context (low latency) | 14.0 MB | CC-BY-SA-4.0 |
| oido-es-ctc-small-int8 | Spanish | streaming + full-context in one file | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 |
Streaming-capable files stream automatically in the firmware's microphone mode and in live_demo.py.
Use
git clone https://github.com/lokutor-ai/oido && cd oido
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo4.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone
How it was made
We took NVIDIA's stt_en_conformer_ctc_small (CC-BY-4.0)
and fine-tuned it with 4-bit quantization-aware training for 8,000 steps. Linear layers inside the Conformer blocks use
int4 per-channel weights; the front end and output head use int8; activations and attention are int8. The training
used public corpora:
- Common Voice 17 and VoxPopuli (CC0);
- LibriSpeech, MLS English, AMI and VCTK (CC-BY-4.0);
- People's Speech, with transcripts regenerated by NVIDIA parakeet-tdt-0.6b-v2;
- OpenSLR 70 and 83 (CC-BY-SA-4.0).
License
This model is released under CC-BY-SA-4.0: it is derived from NVIDIA's CC-BY-4.0 model and trained on data that includes share-alike sources. The Oído engine and firmware are GPLv3, with commercial licenses available from Lokutor.
Model tree for lokutor-ai/oido-ctc-small-int4
Base model
nvidia/stt_en_conformer_ctc_smallDatasets used to train lokutor-ai/oido-ctc-small-int4
MLCommons/peoples_speech
facebook/voxpopuli
Evaluation results
- WER (on-chip int4 arithmetic, greedy) on LibriSpeech (clean)test set self-reported4.610
- WER (on-chip int4 arithmetic, greedy) on LibriSpeech (other)test set self-reported9.980