--- license: cc-by-4.0 language: - es pipeline_tag: automatic-speech-recognition base_model: nvidia/stt_en_conformer_ctc_small tags: - esp32 - esp32-s3 - microcontroller - tinyml - edge-ai - on-device - int8 - conformer - ctc - streaming - spanish datasets: - mozilla-foundation/common_voice_17_0 - facebook/voxpopuli - facebook/multilingual_librispeech - google/fleurs model-index: - name: oido-es-ctc-small-int8 results: - task: type: automatic-speech-recognition name: Speech Recognition dataset: name: Common Voice 17 (es) type: mozilla-foundation/common_voice_17_0 config: es split: test metrics: - type: wer value: 19.99 name: WER (on-chip int8 arithmetic, greedy) - type: wer value: 13.75 name: WER (on-chip int8 arithmetic, with language model) - task: type: automatic-speech-recognition name: Speech Recognition dataset: name: Multilingual LibriSpeech (es) type: facebook/multilingual_librispeech config: spanish split: test metrics: - type: wer value: 14.53 name: WER (on-chip int8 arithmetic, greedy) - type: wer value: 10.90 name: WER (on-chip int8 arithmetic, with language model) - task: type: automatic-speech-recognition name: Speech Recognition dataset: name: VoxPopuli (es) type: facebook/voxpopuli config: es split: test metrics: - type: wer value: 20.25 name: WER (on-chip int8 arithmetic, greedy) - type: wer value: 15.68 name: WER (on-chip int8 arithmetic, with language model) - task: type: automatic-speech-recognition name: Speech Recognition dataset: name: FLEURS (es_419) type: google/fleurs config: es_419 split: test metrics: - type: wer value: 16.62 name: WER (on-chip int8 arithmetic, greedy) - type: wer value: 11.32 name: WER (on-chip int8 arithmetic, with language model) --- # Oído en español: reconocimiento de voz en un microcontrolador de 5 dólares *¡Oído!* is what cooks call out in a Spanish kitchen to confirm an order: *heard, got it*. Spanish speech-to-text that runs entirely on an **ESP32-S3** (240 MHz dual core, 8 MB PSRAM, 16 MB flash, no neural accelerator): any Spanish sentence, no cloud, no command list. It is NVIDIA's 13 M-parameter [`stt_en_conformer_ctc_small`](https://huggingface.co/nvidia/stt_en_conformer_ctc_small) with a new Spanish vocabulary, fine-tuned on 2,492 hours of Spanish (Common Voice, VoxPopuli, Multilingual LibriSpeech, FLEURS) with noise, music, babble and reverberation augmentation, quantized to int8, plus a 1.3 M-parameter Spanish language model for on-chip beam search. **Code, firmware and tools:** [github.com/lokutor-ai/oido](https://github.com/lokutor-ai/oido) (GPLv3, commercial licenses available). English models: [int8](https://huggingface.co/lokutor-ai/oido-ctc-small-int8), [streaming](https://huggingface.co/lokutor-ai/oido-ctc-small-stream-int8). Word error rate (%) on the same 400 evenly spaced utterances per test set (333 for FLEURS after dropping references with digits), same normalization for every system (lowercase, accents kept, punctuation removed): | System | Runs on | Common Voice | MLS | VoxPopuli | FLEURS | |---|---|---|---|---|---| | **This model + language model** | ESP32-S3 | **13.3** | **10.6** | **15.5** | **11.3** | | This model, greedy | ESP32-S3 | 20.3 | 14.2 | 20.0 | 16.9 | | This model, streaming (32-frame chunks) + language model | ESP32-S3 | 15.2 | 12.1 | 16.3 | 12.8 | | Whisper tiny (multilingual), fp32 | laptop | 33.0 | 21.5 | 28.7 | 17.1 | On the complete test sets (60 hours): 13.75 / 10.90 / 15.68 / 11.32 with the language model, 19.99 / 14.53 / 20.25 / 16.62 greedy. **Read this fairly.** We fine-tuned on the training splits of these corpora and Whisper tiny is zero-shot, so the comparison favors us on these domains; expect higher error on phone calls, strong regional accents and specialized vocabulary. There is no Spanish noise benchmark yet. The language model weights (0.5 / 1.5, stored in `oido_es.tlm`) were chosen from a small grid on subsets of the test sets, on a flat optimum. The model writes numbers as words. **Status (3 October 2026).** WERs come from the host build of the firmware engine (same C code, same int8 arithmetic). Under Espressif's QEMU emulator the firmware's transcripts match it on most utterances; the last bit of the math library can change a word on uncertain ones. Speed is estimated (about 0.7–0.95× real time in utterance mode) from exact instruction counts and has not been measured on a board yet. ## Which Oído model? One model per language; the models that stream also run in full-context (utterance) mode. | Model | Language | Modes | Size | License | |---|---|---|---|---| | [oido-ctc-small-int8](https://huggingface.co/lokutor-ai/oido-ctc-small-int8) | English | utterance only (best accuracy) | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 | | [oido-ctc-small-int4](https://huggingface.co/lokutor-ai/oido-ctc-small-int4) | English | utterance only (smallest) | 8.3 MB | CC-BY-SA-4.0 | | [oido-ctc-small-stream-int8](https://huggingface.co/lokutor-ai/oido-ctc-small-stream-int8) | English | **streaming** + full-context (low latency) | 14.0 MB | CC-BY-SA-4.0 | | [oido-es-ctc-small-int8](https://huggingface.co/lokutor-ai/oido-es-ctc-small-int8) | Spanish | **streaming + full-context** in one file | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 | Streaming-capable files stream automatically in the firmware's microphone mode and in `live_demo.py`. ## Use ```bash git clone https://github.com/lokutor-ai/oido && cd oido cd esp32/host && make && python live_demo.py --model es # laptop microphone, streams while you speak ../tools/flash.sh /dev/ttyUSB0 ../../models/oido_es.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone ``` ## Files - `oido_es.tnm`: int8 weights in the TNM1 format (14.0 MB), produced by `train/export_nemo.py`; the Spanish vocabulary is embedded. It works in utterance mode and in streaming mode (chunk 32 frames, left context 128). - `oido_es.tlm`: the Spanish GRU language model (1.3 MB), with its recommended beam-search weights in the header. - `tokenizer.model`: the Spanish SentencePiece vocabulary (1024 unigram pieces). ## License and attribution Derived from NVIDIA's `stt_en_conformer_ctc_small`, licensed CC-BY-4.0. Changes: new Spanish vocabulary, fine-tuned for Spanish, noise robustness and streaming, quantized to int8 and repacked. Fine-tuning data is CC0 or CC-BY (Common Voice 17, VoxPopuli, Multilingual LibriSpeech, FLEURS; MUSAN for augmentation), and unvalidated Common Voice clips were kept only where NVIDIA's `stt_es_conformer_ctc_large` (CC-BY-4.0) agreed with the reference, so this model and its language model are released under **CC-BY-4.0**. The engine and firmware are GPLv3, with commercial licenses from [Lokutor](https://lokutor.com).