Laurin-myreha commited on
Commit
bde49f6
·
verified ·
1 Parent(s): 5ae15b3

model card = package README; error-tail figure

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +217 -76
  3. docs/buckeye_ccdf.png +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ docs/buckeye_ccdf.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -3,121 +3,262 @@ license: other
3
  license_name: nyra-health-non-commercial-research-license
4
  license_link: LICENSE.md
5
  language: [en]
6
- library_name: nyra-align
7
  base_model: microsoft/wavlm-large
8
  tags: [forced-alignment, speech, phonetics, wavlm, verbatim, disfluency, timestamps]
9
  ---
10
 
11
- # nyra_forced_aligner_en
12
 
13
- Verbatim-aware, noise-robust **forced aligner** for English: word and phone
14
- timestamps for conversational speech, including fillers (`[UH]`, `[UM]`),
15
- laughter, breaths and word cut-offs (`wor-*`). Runs **without Kaldi** through the
16
- [`nyra-forced-aligner`](https://github.com/nyrahealth/nyra-forced-aligner) package.
17
 
18
- This is the **standard** model, for clean and lightly noisy audio. A Pro model, more robust to noise and adverse acoustic conditions, is available on request.
 
19
 
20
- **This repository is self-contained**: it bundles the acoustic model (`model.npz`,
21
- `meta.json`, `lexicon.txt`, `phone_map.json`, ~3 MB) *and* the WavLM-large
22
- front-end weights (`wavlm/`), so one download gives you everything.
 
23
 
24
- ## Usage
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
 
26
  ```python
27
  from nyra_align import Aligner
28
 
29
- aligner = Aligner() # this model is the default (downloads + caches)
30
- # aligner = Aligner("nyralabs/nyra_forced_aligner_en") # equivalent, explicit repo id
31
- result = aligner.align("interview.wav", "so i uhm [laughter] i went there")
 
 
32
 
33
  for w in result.words:
34
- print(f"{w.word:>12s} {w.start:7.3f} {w.end:7.3f} {w.type}") # type: w/f/s/c
35
- result.to_textgrid("interview.TextGrid") # Praat
 
36
  result.to_json("interview.json")
37
  ```
38
 
39
- While this repository is private, set `HF_TOKEN` (or `huggingface-cli login`) with access to the `nyralabs` org.
 
40
 
41
- Command line (single file, or an MFA-style corpus of paired `.wav` + `.txt`/`.lab`):
 
42
 
43
  ```bash
44
- nyra-align audio.wav --text "full transcript ..." --out out.TextGrid
45
- nyra-align corpus_dir/ --out-dir aligned/ --format textgrid
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
46
  ```
47
 
48
- Transcripts may contain verbatim event tokens; out-of-lexicon words are
49
- phonemized with espeak-ng (`phonemizer`) through the same mapping used at
50
- training time. Long recordings are handled in one pass (8 min of audio in
51
- ~3 s on one GPU; CPU works but is ~100x slower).
 
 
52
 
53
- ## Model
54
 
55
- WavLM-large layer 23 (20 ms frames, upsampled to 10 ms) -> per-utterance CMVN ->
56
- supervised LDA applied **directly** to the 1024-dim WavLM features (no PCA, no
57
- frame splicing, 1024 -> 40) -> global MLLT -> triphone GMM-HMM (2500 tied states,
58
- 6000 Gaussians, 55 phone classes = 39 speech phones + 15 vocal-event units + SIL).
59
- Trained on ~150 h of verbatim-transcribed conversational English.
60
 
61
- The LDA projection is fitted on clean speech.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
 
63
- ## Results (word boundaries, mean of |onset| and |offset| error)
 
64
 
65
- Full test sets; DP-matched canonical words; laughter tokens excluded from the
66
- word metric for every system; 100% of utterances aligned.
67
 
68
- | Test set | MAE | MedAE | mean IoU | acc <= 25 ms | acc <= 50 ms |
69
- |---|---:|---:|---:|---:|---:|
70
- | Buckeye (conversational, 80,475 words) | **19.9** | 13.9 | 0.830 | 79.5 | 93.7 |
71
- | TIMIT (read, 14,038 words) | 22.5 | 15.0 | 0.838 | 75.7 | 92.1 |
72
- | FluencyBank (stuttered, 5,487 words) | 77.1 | 25.0 | 0.743 | 58.0 | 82.4 |
73
 
74
- Buckeye under added noise (MAE in ms; MUSAN = real noise recordings, Gauss = white noise):
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
75
 
76
- | SNR | MUSAN 20 | 10 | 5 | 0 dB | Gauss 20 | 10 | 5 | 0 dB |
77
- |---|---:|---:|---:|---:|---:|---:|---:|---:|
78
- | this model | 20.0 | 22.2 | 26.9 | 56.3 | 20.2 | 25.3 | 38.9 | 110.9 |
79
- | Montreal Forced Aligner `english_us_arpa` (wide beam) | 33.3 | 65.3 | 127.2 | 281.2 | 36.2 | 74.4 | 156.3 | 416.9 |
 
 
 
 
 
 
80
 
81
- For reference, MFA `english_us_arpa` scores 25.3 ms MAE on clean Buckeye with a
82
- widened search beam (49.0 ms at its shipped default beam).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
83
 
84
  `nyra_forced_aligner_en_pro` delivers roughly the same accuracy on clean speech
85
- and is meaningfully more robust to noise and adverse acoustic conditions; it is
86
- not publicly released and available on request.
87
 
88
- ## Limitations
 
89
 
90
- - English only; the phone inventory and lexicon are espeak-based.
91
- - 10 ms frame grid; median boundary error is ~14 ms on conversational speech.
92
- - Like other GMM-HMM aligners, word boundaries adjacent to long pauses tend to
93
- extend ~25 ms into the pause.
94
- - The transcript must be verbatim for best results; missing or extra words are
95
- absorbed into neighbouring words.
96
 
97
- ## Files
 
 
 
 
 
 
 
 
98
 
99
- | file | content |
100
- |---|---|
101
- | `model.npz` | GMM parameters, transition model, triphone->pdf table, LDA/MLLT matrices |
102
- | `meta.json` | topology, phone table, feature chain, decoding config |
103
- | `lexicon.txt` | 41k-word pronunciation lexicon |
104
- | `phone_map.json` | espeak IPA -> phone mapping used for out-of-lexicon words |
105
- | `wavlm/` | WavLM-large weights (Microsoft; redistributed from `microsoft/wavlm-large`) |
106
 
107
- ## License
108
 
109
- The model files in this repository and **the Outputs they generate** (word and
110
- phone timestamps, alignments, segmentations) are released under the
111
- [nyra health Non-Commercial Research License](LICENSE.md): free for research
112
- and other non-commercial use. **Any commercial use of the model or of its
113
- Outputs — including using alignments or timestamps to train, fine-tune or
114
- improve models intended for commercial use — requires a commercial license**
115
- from nyra health GmbH (licensing@nyra-labs.com).
116
 
117
- The inference code (`nyra-forced-aligner`) is MIT-licensed. The `wavlm/`
118
- directory contains Microsoft's WavLM-large weights, redistributed unchanged
119
- under their own license.
 
 
 
 
 
 
 
120
 
121
- ## Citation
 
 
 
 
 
 
 
122
 
123
- Paper under review (ICASSP 2026). Until then please cite this repository.
 
 
3
  license_name: nyra-health-non-commercial-research-license
4
  license_link: LICENSE.md
5
  language: [en]
6
+ library_name: nyra-forced-aligner
7
  base_model: microsoft/wavlm-large
8
  tags: [forced-alignment, speech, phonetics, wavlm, verbatim, disfluency, timestamps]
9
  ---
10
 
11
+ # nyra-forced-aligner
12
 
13
+ [![Models](https://img.shields.io/badge/%F0%9F%A4%97%20models-nyralabs-yellow)](https://huggingface.co/nyralabs)
 
 
 
14
 
15
+ **The most accurate word timing you can get for conversational speech:
16
+ verbatim-aware, noise-robust, and fast.**
17
 
18
+ [Blog post](https://nyra-labs.com/research/nyra-forced-aligner) ·
19
+ [Paper](#) <!-- TODO: ICASSP 2026 link --> ·
20
+ [Models](https://huggingface.co/nyralabs) ·
21
+ [CrisperWhisper](https://github.com/nyrahealth/CrisperWhisper)
22
 
23
+ A forced aligner answers one question: given a recording and its transcript,
24
+ *when* was each word said? Most aligners were built for read speech and clean
25
+ audio, and they quietly fall apart on the material people actually need to
26
+ align: conversations full of `um`s, laughter, restarts and half-finished words,
27
+ recorded in rooms that are not studios.
28
+
29
+ `nyra-forced-aligner` was built for exactly that material.
30
+
31
+ - **Vocal sounds are alignable.** Fillers (`[UH]`, `[UM]`), laughter, breaths,
32
+ coughs and word cut-offs (`th-`) are units of the model like any phone,
33
+ so they get timestamps of their own instead of being smeared into the
34
+ neighbouring words.
35
+ - **Made for CrisperWhisper transcripts.** Trained on the verbatim
36
+ transcription convention of
37
+ [CrisperWhisper 2.0](https://github.com/nyrahealth/CrisperWhisper): feed
38
+ its output straight in and every filler, repetition and event token is
39
+ aligned.
40
+ - **Accurate where it matters.** 19.9 ms mean word-boundary error on
41
+ conversational English (Buckeye): 21% lower than the best configuration of
42
+ the Montreal Forced Aligner, half the error of the best ASR-with-timestamps
43
+ system, and the only system whose error tail drops below 1% of words before
44
+ 100 ms.
45
+ - **Noise-robust.** Every utterance stays aligned down to 0 dB SNR, where
46
+ classical aligners lose or mangle most of them. A Pro model goes further for
47
+ difficult recordings.
48
+ - **Fast and longform.** Batched WavLM features, GPU emissions and a compiled
49
+ beam Viterbi: eight minutes of audio align in about three seconds on one
50
+ GPU.
51
+
52
+ ## Install
53
+
54
+ ```bash
55
+ pip install nyra-forced-aligner
56
+
57
+ # espeak-ng is needed to pronounce words outside the built-in 41k-word lexicon:
58
+ # apt install espeak-ng (macOS: brew install espeak-ng)
59
+ ```
60
+
61
+ Runs on CPU, but a CUDA GPU is strongly recommended (WavLM-large front-end).
62
+ The model is downloaded from the Hugging Face Hub on first use.
63
+
64
+ ## Quickstart
65
 
66
  ```python
67
  from nyra_align import Aligner
68
 
69
+ aligner = Aligner() # nyralabs/nyra_forced_aligner_en
70
+ # aligner = Aligner(pro=True) # noise-robust model
71
+
72
+ result = aligner.align("interview.wav",
73
+ "so i [UM] i went there on on thursday [laughter]")
74
 
75
  for w in result.words:
76
+ print(f"{w.start:7.3f} {w.end:7.3f} {w.word:<12s} {w.type}") # w / f(iller) / s(ound) / c(ut-off)
77
+
78
+ result.to_textgrid("interview.TextGrid") # Praat
79
  result.to_json("interview.json")
80
  ```
81
 
82
+ Any audio length works. For the best timings follow the
83
+ [transcript conventions](#transcript-conventions) below.
84
 
85
+ Command line, single file or a whole corpus of paired `.wav` + `.txt`/`.lab`
86
+ files (MFA-style layout):
87
 
88
  ```bash
89
+ nyra-align interview.wav --text-file interview.txt --out interview.TextGrid
90
+ nyra-align corpus_dir/ --pro --out-dir aligned/ --format textgrid
91
+ ```
92
+
93
+ ### Transcript conventions
94
+
95
+ The aligner can only place what the transcript contains, so for the best
96
+ timings the transcript should be **verbatim: write what you hear.**
97
+
98
+ - **Cut-offs.** Interrupted words and word fragments are marked with a
99
+ trailing hyphen: `th-`, `w-`, `resched-`.
100
+ - **Fillers.** Filled pauses are bracketed and upper-cased: `[UH]`, `[UM]`.
101
+ - **Vocal sound events.** Non-speech sounds are bracketed: `[laughter]`,
102
+ `[cough]`, `[breath]`, `[sigh]`, `[sniff]`, `[lipsmack]`,
103
+ `[throatclearing]`, `[yawn]`, `[noise]`.
104
+ - **Numbers, dates, times and emails** are written out exactly as spoken
105
+ (`March third at nine thirty`), not as digits or symbols.
106
+ - **Everything else** is transcribed word for word. Normal casing and
107
+ punctuation are fine (the aligner ignores both); repetitions, false starts,
108
+ fillers, fragments and sounds are all left in place.
109
+
110
+ ```text
111
+ so we we need to, to reschedule the th- Thursday meeting to [UH] March third at nine thirty [laughter]
112
  ```
113
 
114
+ What happens when a transcript deviates: a word that was spoken but is missing
115
+ from the transcript gets absorbed into its neighbours, and a word in the
116
+ transcript that was never spoken is squeezed into a few frames, so both shift
117
+ the surrounding boundaries. Event tags are matched in any casing; a tag the
118
+ model does not know (`[music]`) is left out and listed in `result.skipped`,
119
+ as is any token that cannot be pronounced.
120
 
121
+ ### With CrisperWhisper 2.0
122
 
123
+ [CrisperWhisper 2.0](https://github.com/nyrahealth/CrisperWhisper) produces
124
+ transcripts in exactly this convention, so its verbatim output can be aligned
125
+ as is:
 
 
126
 
127
+ ```python
128
+ from crisperwhisper import CrisperWhisperModel
129
+ from nyra_align import Aligner
130
+
131
+ audio = "interview.wav"
132
+ transcript = CrisperWhisperModel("large").transcribe(audio, language="en").text
133
+ result = Aligner().align(audio, transcript)
134
+
135
+ for w in result.words:
136
+ print(f"{w.start:7.3f} {w.end:7.3f} {w.word}")
137
+ result.to_textgrid("interview.TextGrid")
138
+ ```
139
+
140
+ ### Models
141
+
142
+ | Shorthand | Hugging Face ID | Use it for |
143
+ |-----------|-----------------|------------|
144
+ | `Aligner()` (default) | `nyralabs/nyra_forced_aligner_en` | Clean and lightly noisy audio; most accurate on clean speech |
145
+ | `Aligner(pro=True)` | `nyralabs/nyra_forced_aligner_en_pro` | **Pro**: roughly the same accuracy on clean speech, meaningfully more robust to noise and adverse acoustic conditions. Not publicly released; available on request |
146
 
147
+ Only English is available at the moment; `Aligner(language="de")` raises
148
+ `NotImplementedError`.
149
 
150
+ ## Performance
 
151
 
152
+ Mean absolute word-boundary error on **Buckeye** (conversational English,
153
+ 2,008 test recordings, 80,475 words), lower is better. Every system is scored
154
+ with the same protocol: DP word matching on canonicalized forms, laughter
155
+ tokens excluded, MAE = mean over words of (|onset error| + |offset error|)/2.
 
156
 
157
+ | # | System | Type | MAE (ms) | MedAE | ≤25 ms | ≤50 ms | ≤100 ms | mIoU |
158
+ |--:|--------|------|---------:|------:|-------:|-------:|--------:|-----:|
159
+ | 1 | **nyra_forced_aligner_en** | forced aligner | **19.9** | 13.9 | **79.5%** | **93.7%** | **98.3%** | **0.830** |
160
+ | 2 | Montreal Forced Aligner `english_us_arpa`, wide beam¹ | forced aligner | 25.3 | 13.0 | 76.7% | 92.2% | 97.3% | 0.826 |
161
+ | 3 | Montreal Forced Aligner `english_mfa`, wide beam¹ | forced aligner | 35.6 | 15.0 | 71.3% | 87.5% | 95.2% | 0.790 |
162
+ | 4 | CrisperWhisper 2.0 | ASR | 38.7 | 25.0 | 54.2% | 78.7% | 91.7% | 0.717 |
163
+ | 5 | Qwen3-ForcedAligner-0.6B | forced aligner | 39.5 | 25.0 | 52.8% | 85.9% | 96.6% | 0.711 |
164
+ | 6 | xAI Grok Speech-to-Text | ASR | 45.5 | 34.5 | 58.3% | 82.8% | 94.0% | 0.630 |
165
+ | 7 | Montreal Forced Aligner `english_us_arpa`, default beam | forced aligner | 48.0 | 13.5 | 75.2% | 90.4% | 95.6% | 0.810 |
166
+ | 8 | Montreal Forced Aligner `english_mfa`, default beam | forced aligner | 48.9 | 15.5 | 70.7% | 86.7% | 94.3% | 0.782 |
167
+ | 9 | MMS-FA (torchaudio) | forced aligner | 53.0 | 34.9 | 57.5% | 81.3% | 92.8% | 0.620 |
168
+ | 10 | ElevenLabs Scribe v2 | ASR | 57.8 | 50.0 | 47.6% | 75.8% | 94.8% | 0.540 |
169
+ | 11 | Charsiu | forced aligner | 70.3 | 28.0 | 51.8% | 68.5% | 85.1% | 0.649 |
170
+ | 12 | NeMo Forced Aligner | forced aligner | 72.6 | 62.5 | 26.8% | 49.8% | 79.3% | 0.443 |
171
+ | 13 | WhisperX (wav2vec2 alignment) | forced aligner | 80.9 | 61.5 | 35.1% | 68.1% | 94.1% | 0.456 |
172
+ | 14 | CTC-Segmentation | forced aligner | 81.0 | 36.0 | 41.5% | 69.5% | 89.3% | 0.635 |
173
+ | 15 | Deepgram Nova-3 | ASR | 86.9 | 71.0 | 18.1% | 36.4% | 68.9% | 0.496 |
174
+ | 16 | Cartesia Ink-Whisper | ASR | 124.8 | 46.3 | 36.0% | 60.4% | 82.8% | 0.598 |
175
 
176
+ <sub>**MAE** and **MedAE** are over the per-word error (|onset error| +
177
+ |offset error|)/2. The **≤ t ms** columns follow the convention of Rousso et
178
+ al. (2024): the share of words whose *end* timestamp is within t of the
179
+ reference (one boundary, not the per-word mean). **Forced aligners** are given
180
+ the reference words and only predict their timing; every word is scored. **ASR** systems produce their own words, which
181
+ are text-aligned to the reference, and only matched words are scored
182
+ (88–93% of words), so their timing is judged independently of their
183
+ transcription errors. ¹ MFA's shipped default beam (10) prunes the correct
184
+ path on conversational speech; the "wide beam" rows use `beam 400 / retry
185
+ 4000`, the best configuration we found for it.</sub>
186
 
187
+ The mean hides where the difference really is. Averaged over all words, most
188
+ systems are within a few tens of milliseconds; the gap is in the **tail**, the
189
+ words that are misaligned by 100 ms or more, which is what breaks downstream
190
+ use. Each curve shows the share of words whose per-word error (mean of the
191
+ onset and offset errors, the quantity behind MAE) is larger than *t*; note
192
+ this is not the single-boundary ≤ t ms column of the table:
193
+
194
+ ![Buckeye word-boundary error tails](https://huggingface.co/nyralabs/nyra_forced_aligner_en/resolve/main/docs/buckeye_ccdf.png)
195
+
196
+ ### Under noise
197
+
198
+ The same Buckeye recordings with added noise, at signal-to-noise ratios from
199
+ 20 dB down to 0 dB (MAE in ms). MUSAN is real noise recordings; Gaussian is
200
+ white noise, which the model never saw in training.
201
+
202
+ | System | MUSAN 20 | 10 | 5 | 0 dB | Gauss 20 | 10 | 5 | 0 dB |
203
+ |--------|---:|---:|---:|---:|---:|---:|---:|---:|
204
+ | **nyra_forced_aligner_en** | 20.0 | 22.2 | 26.9 | 56.3 | 20.2 | 25.3 | 38.9 | 110.9 |
205
+ | CrisperWhisper 2.0² | 38.7 | 42.4 | 48.0 | 55.7 | 40.1 | 50.9 | 67.3 | 103.4 |
206
+ | Qwen3-ForcedAligner-0.6B | 40.6 | 54.3 | 78.7 | 153.9 | 43.4 | 85.2 | 153.9 | 339.2 |
207
+ | MMS-FA (torchaudio) | 52.8 | 54.9 | 57.7 | 67.0 | 52.0 | 55.2 | 61.0 | 86.1 |
208
+ | Montreal Forced Aligner `english_us_arpa`, wide beam | 33.3 | 65.3 | 127.2 | 281.2 | 36.2 | 74.4 | 156.3 | 416.9 |
209
+ | WhisperX (wav2vec2 alignment) | 81.7 | 90.4 | 138.1 | 468.5 | 84.2 | 126.3 | 415.2 | 2720.6 |
210
+
211
+ <sub>² CrisperWhisper 2.0 is an ASR system: at 0 dB it still recognises only
212
+ 29% of the words, and only those are scored. All forced aligners above keep
213
+ 100% of words at every SNR.</sub>
214
 
215
  `nyra_forced_aligner_en_pro` delivers roughly the same accuracy on clean speech
216
+ and is meaningfully more robust to noise and adverse acoustic conditions.
 
217
 
218
+ Results on read speech (TIMIT) and stuttered speech (FluencyBank), the full
219
+ set of ablations, and how the models are trained are in the paper.
220
 
221
+ ## What else is in the box
 
 
 
 
 
222
 
223
+ | Option | What it does |
224
+ |--------|--------------|
225
+ | `Aligner(pro=True)` | Noise-robust model |
226
+ | `Aligner(device="cuda" / "cpu")` | Device selection (auto by default) |
227
+ | `Aligner(precision="fp32")` | WavLM in fp32 instead of fp16; bit-exact parity with the reference implementation (fp16 moves ≤30 ms on ~0.5% of word boundaries) |
228
+ | `Aligner(wavlm_batch=16)` | Batch size for the 30 s WavLM windows on long audio (tune to VRAM) |
229
+ | `aligner.align_features(feats, text)` | Re-align a transcript against cached features (decode is ~6 ms per 3.5 s utterance) |
230
+ | `result.events` | Just the filler / laughter / cut-off tokens |
231
+ | `result.silences` | Silence segments between words |
232
 
233
+ ## How it works
 
 
 
 
 
 
234
 
235
+ The full story is in the deep-dive post:
236
 
237
+ - [The nyra forced aligner: precise word timing for real conversational speech](https://nyra-labs.com/research/nyra-forced-aligner).
238
+ Why classical aligners break on spontaneous and noisy speech, how frozen
239
+ WavLM features, vocal-event units and a noise-pooled projection fix it, and
240
+ how the model bootstraps itself from verbatim transcripts without any
241
+ manual boundary labels.
 
 
242
 
243
+ In short: WavLM-large hidden states are projected by a supervised LDA onto 40
244
+ dimensions and modelled by a triphone GMM-HMM whose inventory holds 15
245
+ vocal-event units next to the 39 English phones. At inference the package
246
+ builds the alignment graph for the transcript (optional silence between words,
247
+ full context across word boundaries) and runs a compiled beam Viterbi over the
248
+ model's likelihoods, with a beam ladder so that no utterance is ever silently
249
+ dropped. Words outside the built-in lexicon are phonemized with espeak through
250
+ the same mapping used at training time.
251
+
252
+ ## License
253
 
254
+ The model files are released under the
255
+ [nyra health Non-Commercial Research License](https://huggingface.co/nyralabs/nyra_forced_aligner_en/blob/main/LICENSE.md):
256
+ free for research and other non-commercial use. This license covers the
257
+ model **and the alignments it produces**: any commercial use of the
258
+ timestamps, and any use of them to train, fine-tune or otherwise improve
259
+ models intended for commercial use, requires a commercial license. The Pro
260
+ model is available under commercial license only. For commercial licensing
261
+ of either, [contact nyra health](mailto:licensing@nyra-labs.com).
262
 
263
+ The model repositories bundle Microsoft's WavLM-large weights, redistributed
264
+ unchanged under their upstream license.
docs/buckeye_ccdf.png ADDED

Git LFS Details

  • SHA256: 6aec1125d2ff7e2df522d818a1c7b329959dd149a9d4a9391379ffcc4b10c395
  • Pointer size: 131 Bytes
  • Size of remote file: 284 kB