danivs10 commited on
Commit
f9468e4
·
verified ·
1 Parent(s): 75c1636

Release int4 Oído model (8.3 MB)

Browse files
Files changed (4) hide show
  1. .gitattributes +1 -0
  2. README.md +91 -0
  3. nemo4.tnm +3 -0
  4. tokenizer.model +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ nemo4.tnm filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,91 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-sa-4.0
3
+ language:
4
+ - en
5
+ pipeline_tag: automatic-speech-recognition
6
+ base_model: nvidia/stt_en_conformer_ctc_small
7
+ tags:
8
+ - esp32
9
+ - esp32-s3
10
+ - microcontroller
11
+ - tinyml
12
+ - edge-ai
13
+ - on-device
14
+ - int4
15
+ - quantization-aware-training
16
+ - conformer
17
+ - ctc
18
+ datasets:
19
+ - openslr/librispeech_asr
20
+ - mozilla-foundation/common_voice_17_0
21
+ - facebook/voxpopuli
22
+ - MLCommons/peoples_speech
23
+ model-index:
24
+ - name: oido-ctc-small-int4
25
+ results:
26
+ - task:
27
+ type: automatic-speech-recognition
28
+ name: Speech Recognition
29
+ dataset:
30
+ name: LibriSpeech (clean)
31
+ type: openslr/librispeech_asr
32
+ config: clean
33
+ split: test
34
+ metrics:
35
+ - type: wer
36
+ value: 4.61
37
+ name: WER (on-chip int4 arithmetic, greedy)
38
+ - task:
39
+ type: automatic-speech-recognition
40
+ name: Speech Recognition
41
+ dataset:
42
+ name: LibriSpeech (other)
43
+ type: openslr/librispeech_asr
44
+ config: other
45
+ split: test
46
+ metrics:
47
+ - type: wer
48
+ value: 9.98
49
+ name: WER (on-chip int4 arithmetic, greedy)
50
+ ---
51
+
52
+ # Oído int4: 8.3 MB speech recognition for the ESP32-S3
53
+
54
+ *¡Oído!* is Spanish kitchen slang for *heard, got it*.
55
+
56
+ This is the compact profile of [Oído](https://github.com/lokutor-ai/oido): open-vocabulary English speech recognition
57
+ that runs entirely on an ESP32-S3, with no cloud and no NPU. At **8.3 MB** it leaves a 6 MB app partition free on a 16 MB
58
+ flash module for your own application code (`esp32/firmware/partitions_nemo4.csv`).
59
+
60
+ | LibriSpeech WER (%) | test-clean | test-other | Size |
61
+ |---|---|---|---|
62
+ | **This model** (int4, greedy, on-chip arithmetic) | **4.61** | **9.98** | 8.3 MB |
63
+ | [Oído int8](https://huggingface.co/lokutor-ai/oido-ctc-small-int8) | 3.70 | 8.23 | 14.0 MB |
64
+ | Espressif MultiNet7 on the same chip (ESP-SR benchmark) | 8.5 | 21.3 | 2.9 MB |
65
+
66
+ The int4 weights also run about 10% faster than int8 (estimated RTF 0.68–0.82 from exact QEMU instruction counts; not
67
+ yet measured on silicon).
68
+
69
+ ## Use
70
+
71
+ ```bash
72
+ git clone https://github.com/lokutor-ai/oido && cd oido
73
+ esp32/tools/flash.sh /dev/ttyUSB0 models/nemo4.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone
74
+ ```
75
+
76
+ ## How it was made
77
+
78
+ We took NVIDIA's [`stt_en_conformer_ctc_small`](https://huggingface.co/nvidia/stt_en_conformer_ctc_small) (CC-BY-4.0)
79
+ and fine-tuned it with 4-bit quantization-aware training for 8,000 steps. Linear layers inside the Conformer blocks use
80
+ int4 per-channel weights; the front end and output head use int8; activations and attention are int8. The training
81
+ used public corpora:
82
+ - Common Voice 17 and VoxPopuli (CC0);
83
+ - LibriSpeech, MLS English, AMI and VCTK (CC-BY-4.0);
84
+ - People's Speech, with transcripts regenerated by NVIDIA parakeet-tdt-0.6b-v2;
85
+ - OpenSLR 70 and 83 (CC-BY-SA-4.0).
86
+
87
+ ## License
88
+
89
+ This model is released under CC-BY-SA-4.0: it is derived from NVIDIA's CC-BY-4.0 model and trained on data that
90
+ includes share-alike sources. The Oído engine and firmware are GPLv3, with commercial licenses available from
91
+ [Lokutor](https://lokutor.com).
nemo4.tnm ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e6fad1c3a1c0065add9fe3c9f10d6fa075b09b27c11f1a830b35f31fabf7d5e9
3
+ size 8283037
tokenizer.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d9b04033136c5d0413047fe94d2f0ab6cb088d292014bb076ee1700bdac545a9
3
+ size 260411