--- license: cc-by-sa-4.0 language: - en pipeline_tag: automatic-speech-recognition base_model: nvidia/stt_en_conformer_ctc_small tags: - esp32 - esp32-s3 - microcontroller - tinyml - edge-ai - on-device - int4 - quantization-aware-training - conformer - ctc datasets: - openslr/librispeech_asr - mozilla-foundation/common_voice_17_0 - facebook/voxpopuli - MLCommons/peoples_speech model-index: - name: oido-ctc-small-int4 results: - task: type: automatic-speech-recognition name: Speech Recognition dataset: name: LibriSpeech (clean) type: openslr/librispeech_asr config: clean split: test metrics: - type: wer value: 4.61 name: WER (on-chip int4 arithmetic, greedy) - task: type: automatic-speech-recognition name: Speech Recognition dataset: name: LibriSpeech (other) type: openslr/librispeech_asr config: other split: test metrics: - type: wer value: 9.98 name: WER (on-chip int4 arithmetic, greedy) --- # Oído int4: 8.3 MB speech recognition for the ESP32-S3 *Animated demo: the transcripts are Oído's output (sped up). Footage from a physical board is coming.* *¡Oído!* is Spanish kitchen slang for *heard, got it*. This is the compact profile of [Oído](https://github.com/lokutor-ai/oido): open-vocabulary English speech recognition that runs entirely on an ESP32-S3, with no cloud and no NPU. At **8.3 MB** it leaves a 6 MB app partition free on a 16 MB flash module for your own application code (`esp32/firmware/partitions_nemo4.csv`). | LibriSpeech WER (%) | test-clean | test-other | Size | |---|---|---|---| | **This model** (int4, greedy, on-chip arithmetic) | **4.61** | **9.98** | 8.3 MB | | [Oído int8](https://huggingface.co/lokutor-ai/oido-ctc-small-int8) | 3.70 | 8.23 | 14.0 MB | | Espressif MultiNet7 on the same chip (ESP-SR benchmark) | 8.5 | 21.3 | 2.9 MB | The int4 weights also run about 10% faster than int8 (estimated RTF 0.68–0.82 from exact QEMU instruction counts; not yet measured on silicon). **Need lower latency?** The [streaming variant](https://huggingface.co/lokutor-ai/oido-ctc-small-stream-int8) shows text while you speak and delivers the final text about 1.1–1.4 s after you stop (estimated), at some cost in accuracy. ## Which Oído model? One model per language; the models that stream also run in full-context (utterance) mode. | Model | Language | Modes | Size | License | |---|---|---|---|---| | [oido-ctc-small-int8](https://huggingface.co/lokutor-ai/oido-ctc-small-int8) | English | utterance only (best accuracy) | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 | | [oido-ctc-small-int4](https://huggingface.co/lokutor-ai/oido-ctc-small-int4) | English | utterance only (smallest) | 8.3 MB | CC-BY-SA-4.0 | | [oido-ctc-small-stream-int8](https://huggingface.co/lokutor-ai/oido-ctc-small-stream-int8) | English | **streaming** + full-context (low latency) | 14.0 MB | CC-BY-SA-4.0 | | [oido-es-ctc-small-int8](https://huggingface.co/lokutor-ai/oido-es-ctc-small-int8) | Spanish | **streaming + full-context** in one file | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 | Streaming-capable files stream automatically in the firmware's microphone mode and in `live_demo.py`. ## Use ```bash git clone https://github.com/lokutor-ai/oido && cd oido esp32/tools/flash.sh /dev/ttyUSB0 models/nemo4.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone ``` ## How it was made We took NVIDIA's [`stt_en_conformer_ctc_small`](https://huggingface.co/nvidia/stt_en_conformer_ctc_small) (CC-BY-4.0) and fine-tuned it with 4-bit quantization-aware training for 8,000 steps. Linear layers inside the Conformer blocks use int4 per-channel weights; the front end and output head use int8; activations and attention are int8. The training used public corpora: - Common Voice 17 and VoxPopuli (CC0); - LibriSpeech, MLS English, AMI and VCTK (CC-BY-4.0); - People's Speech, with transcripts regenerated by NVIDIA parakeet-tdt-0.6b-v2; - OpenSLR 70 and 83 (CC-BY-SA-4.0). ## License This model is released under CC-BY-SA-4.0: it is derived from NVIDIA's CC-BY-4.0 model and trained on data that includes share-alike sources. The Oído engine and firmware are GPLv3, with commercial licenses available from [Lokutor](https://lokutor.com).