Oído en español: reconocimiento de voz en un microcontrolador de 5 dólares

¡Oído! is what cooks call out in a Spanish kitchen to confirm an order: heard, got it.

Spanish speech-to-text that runs entirely on an ESP32-S3 (240 MHz dual core, 8 MB PSRAM, 16 MB flash, no neural accelerator): any Spanish sentence, no cloud, no command list. It is NVIDIA's 13 M-parameter stt_en_conformer_ctc_small with a new Spanish vocabulary, fine-tuned on 2,492 hours of Spanish (Common Voice, VoxPopuli, Multilingual LibriSpeech, FLEURS) with noise, music, babble and reverberation augmentation, quantized to int8, plus a 1.3 M-parameter Spanish language model for on-chip beam search.

Code, firmware and tools: github.com/lokutor-ai/oido (GPLv3, commercial licenses available). English models: int8, streaming.

Word error rate (%) on the same 400 evenly spaced utterances per test set (333 for FLEURS after dropping references with digits), same normalization for every system (lowercase, accents kept, punctuation removed):

System Runs on Common Voice MLS VoxPopuli FLEURS
This model + language model ESP32-S3 13.3 10.6 15.5 11.3
This model, greedy ESP32-S3 20.3 14.2 20.0 16.9
This model, streaming (32-frame chunks) + language model ESP32-S3 15.2 12.1 16.3 12.8
Whisper tiny (multilingual), fp32 laptop 33.0 21.5 28.7 17.1

On the complete test sets (60 hours): 13.75 / 10.90 / 15.68 / 11.32 with the language model, 19.99 / 14.53 / 20.25 / 16.62 greedy.

Read this fairly. We fine-tuned on the training splits of these corpora and Whisper tiny is zero-shot, so the comparison favors us on these domains; expect higher error on phone calls, strong regional accents and specialized vocabulary. There is no Spanish noise benchmark yet. The language model weights (0.5 / 1.5, stored in oido_es.tlm) were chosen from a small grid on subsets of the test sets, on a flat optimum. The model writes numbers as words.

Status (3 October 2026). WERs come from the host build of the firmware engine (same C code, same int8 arithmetic). Under Espressif's QEMU emulator the firmware's transcripts match it on most utterances; the last bit of the math library can change a word on uncertain ones. Speed is estimated (about 0.7–0.95× real time in utterance mode) from exact instruction counts and has not been measured on a board yet.

Which Oído model?

One model per language; the models that stream also run in full-context (utterance) mode.

Model Language Modes Size License
oido-ctc-small-int8 English utterance only (best accuracy) 14.0 MB (+1.3 MB LM) CC-BY-4.0
oido-ctc-small-int4 English utterance only (smallest) 8.3 MB CC-BY-SA-4.0
oido-ctc-small-stream-int8 English streaming + full-context (low latency) 14.0 MB CC-BY-SA-4.0
oido-es-ctc-small-int8 Spanish streaming + full-context in one file 14.0 MB (+1.3 MB LM) CC-BY-4.0

Streaming-capable files stream automatically in the firmware's microphone mode and in live_demo.py.

Use

git clone https://github.com/lokutor-ai/oido && cd oido
cd esp32/host && make && python live_demo.py --model es            # laptop microphone, streams while you speak
../tools/flash.sh /dev/ttyUSB0 ../../models/oido_es.tnm            # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone

Files

  • oido_es.tnm: int8 weights in the TNM1 format (14.0 MB), produced by train/export_nemo.py; the Spanish vocabulary is embedded. It works in utterance mode and in streaming mode (chunk 32 frames, left context 128).
  • oido_es.tlm: the Spanish GRU language model (1.3 MB), with its recommended beam-search weights in the header.
  • tokenizer.model: the Spanish SentencePiece vocabulary (1024 unigram pieces).

License and attribution

Derived from NVIDIA's stt_en_conformer_ctc_small, licensed CC-BY-4.0. Changes: new Spanish vocabulary, fine-tuned for Spanish, noise robustness and streaming, quantized to int8 and repacked. Fine-tuning data is CC0 or CC-BY (Common Voice 17, VoxPopuli, Multilingual LibriSpeech, FLEURS; MUSAN for augmentation), and unvalidated Common Voice clips were kept only where NVIDIA's stt_es_conformer_ctc_large (CC-BY-4.0) agreed with the reference, so this model and its language model are released under CC-BY-4.0. The engine and firmware are GPLv3, with commercial licenses from Lokutor.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lokutor-ai/oido-es-ctc-small-int8

Finetuned
(5)
this model

Datasets used to train lokutor-ai/oido-es-ctc-small-int8

Evaluation results