Oído en español: reconocimiento de voz en un microcontrolador de 5 dólares
¡Oído! is what cooks call out in a Spanish kitchen to confirm an order: heard, got it.
Spanish speech-to-text that runs entirely on an ESP32-S3 (240 MHz dual core, 8 MB PSRAM, 16 MB flash, no neural
accelerator): any Spanish sentence, no cloud, no command list. It is NVIDIA's 13 M-parameter
stt_en_conformer_ctc_small with a new Spanish vocabulary,
fine-tuned on 2,492 hours of Spanish (Common Voice, VoxPopuli, Multilingual LibriSpeech, FLEURS) with noise, music, babble
and reverberation augmentation, quantized to int8, plus a 1.3 M-parameter Spanish language model for on-chip beam search.
Code, firmware and tools: github.com/lokutor-ai/oido (GPLv3, commercial licenses available). English models: int8, streaming.
Word error rate (%) on the same 400 evenly spaced utterances per test set (333 for FLEURS after dropping references with digits), same normalization for every system (lowercase, accents kept, punctuation removed):
| System | Runs on | Common Voice | MLS | VoxPopuli | FLEURS |
|---|---|---|---|---|---|
| This model + language model | ESP32-S3 | 13.3 | 10.6 | 15.5 | 11.3 |
| This model, greedy | ESP32-S3 | 20.3 | 14.2 | 20.0 | 16.9 |
| This model, streaming (32-frame chunks) + language model | ESP32-S3 | 15.2 | 12.1 | 16.3 | 12.8 |
| Whisper tiny (multilingual), fp32 | laptop | 33.0 | 21.5 | 28.7 | 17.1 |
On the complete test sets (60 hours): 13.75 / 10.90 / 15.68 / 11.32 with the language model, 19.99 / 14.53 / 20.25 / 16.62 greedy.
Read this fairly. We fine-tuned on the training splits of these corpora and Whisper tiny is zero-shot, so the
comparison favors us on these domains; expect higher error on phone calls, strong regional accents and specialized
vocabulary. There is no Spanish noise benchmark yet. The language model weights (0.5 / 1.5, stored in oido_es.tlm) were
chosen from a small grid on subsets of the test sets, on a flat optimum. The model writes numbers as words.
Status (3 October 2026). WERs come from the host build of the firmware engine (same C code, same int8 arithmetic). Under Espressif's QEMU emulator the firmware's transcripts match it on most utterances; the last bit of the math library can change a word on uncertain ones. Speed is estimated (about 0.7–0.95× real time in utterance mode) from exact instruction counts and has not been measured on a board yet.
Which Oído model?
One model per language; the models that stream also run in full-context (utterance) mode.
| Model | Language | Modes | Size | License |
|---|---|---|---|---|
| oido-ctc-small-int8 | English | utterance only (best accuracy) | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 |
| oido-ctc-small-int4 | English | utterance only (smallest) | 8.3 MB | CC-BY-SA-4.0 |
| oido-ctc-small-stream-int8 | English | streaming + full-context (low latency) | 14.0 MB | CC-BY-SA-4.0 |
| oido-es-ctc-small-int8 | Spanish | streaming + full-context in one file | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 |
Streaming-capable files stream automatically in the firmware's microphone mode and in live_demo.py.
Use
git clone https://github.com/lokutor-ai/oido && cd oido
cd esp32/host && make && python live_demo.py --model es # laptop microphone, streams while you speak
../tools/flash.sh /dev/ttyUSB0 ../../models/oido_es.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone
Files
oido_es.tnm: int8 weights in the TNM1 format (14.0 MB), produced bytrain/export_nemo.py; the Spanish vocabulary is embedded. It works in utterance mode and in streaming mode (chunk 32 frames, left context 128).oido_es.tlm: the Spanish GRU language model (1.3 MB), with its recommended beam-search weights in the header.tokenizer.model: the Spanish SentencePiece vocabulary (1024 unigram pieces).
License and attribution
Derived from NVIDIA's stt_en_conformer_ctc_small, licensed CC-BY-4.0. Changes: new Spanish vocabulary, fine-tuned for
Spanish, noise robustness and streaming, quantized to int8 and repacked. Fine-tuning data is CC0 or CC-BY
(Common Voice 17, VoxPopuli, Multilingual LibriSpeech, FLEURS; MUSAN for augmentation), and unvalidated Common Voice clips
were kept only where NVIDIA's stt_es_conformer_ctc_large (CC-BY-4.0) agreed with the reference, so this model and its
language model are released under CC-BY-4.0. The engine and firmware are GPLv3, with commercial licenses from
Lokutor.
Model tree for lokutor-ai/oido-es-ctc-small-int8
Base model
nvidia/stt_en_conformer_ctc_smallDatasets used to train lokutor-ai/oido-es-ctc-small-int8
facebook/multilingual_librispeech
facebook/voxpopuli
Evaluation results
- WER (on-chip int8 arithmetic, greedy) on Common Voice 17 (es)test set self-reported19.990
- WER (on-chip int8 arithmetic, with language model) on Common Voice 17 (es)test set self-reported13.750
- WER (on-chip int8 arithmetic, greedy) on Multilingual LibriSpeech (es)test set self-reported14.530
- WER (on-chip int8 arithmetic, with language model) on Multilingual LibriSpeech (es)test set self-reported10.900
- WER (on-chip int8 arithmetic, greedy) on VoxPopuli (es)test set self-reported20.250
- WER (on-chip int8 arithmetic, with language model) on VoxPopuli (es)test set self-reported15.680
- WER (on-chip int8 arithmetic, greedy) on FLEURS (es_419)test set self-reported16.620
- WER (on-chip int8 arithmetic, with language model) on FLEURS (es_419)test set self-reported11.320