model card = package README; error-tail figure
Browse files- .gitattributes +1 -0
- README.md +217 -76
- docs/buckeye_ccdf.png +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
docs/buckeye_ccdf.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -3,121 +3,262 @@ license: other
|
|
| 3 |
license_name: nyra-health-non-commercial-research-license
|
| 4 |
license_link: LICENSE.md
|
| 5 |
language: [en]
|
| 6 |
-
library_name: nyra-
|
| 7 |
base_model: microsoft/wavlm-large
|
| 8 |
tags: [forced-alignment, speech, phonetics, wavlm, verbatim, disfluency, timestamps]
|
| 9 |
---
|
| 10 |
|
| 11 |
-
#
|
| 12 |
|
| 13 |
-
|
| 14 |
-
timestamps for conversational speech, including fillers (`[UH]`, `[UM]`),
|
| 15 |
-
laughter, breaths and word cut-offs (`wor-*`). Runs **without Kaldi** through the
|
| 16 |
-
[`nyra-forced-aligner`](https://github.com/nyrahealth/nyra-forced-aligner) package.
|
| 17 |
|
| 18 |
-
|
|
|
|
| 19 |
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
|
|
|
| 23 |
|
| 24 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
```python
|
| 27 |
from nyra_align import Aligner
|
| 28 |
|
| 29 |
-
aligner = Aligner()
|
| 30 |
-
# aligner = Aligner(
|
| 31 |
-
|
|
|
|
|
|
|
| 32 |
|
| 33 |
for w in result.words:
|
| 34 |
-
print(f"{w.
|
| 35 |
-
|
|
|
|
| 36 |
result.to_json("interview.json")
|
| 37 |
```
|
| 38 |
|
| 39 |
-
|
|
|
|
| 40 |
|
| 41 |
-
Command line
|
|
|
|
| 42 |
|
| 43 |
```bash
|
| 44 |
-
nyra-align
|
| 45 |
-
nyra-align corpus_dir/ --out-dir aligned/ --format textgrid
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 46 |
```
|
| 47 |
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
|
|
|
|
|
|
| 52 |
|
| 53 |
-
##
|
| 54 |
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
6000 Gaussians, 55 phone classes = 39 speech phones + 15 vocal-event units + SIL).
|
| 59 |
-
Trained on ~150 h of verbatim-transcribed conversational English.
|
| 60 |
|
| 61 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
-
|
|
|
|
| 64 |
|
| 65 |
-
|
| 66 |
-
word metric for every system; 100% of utterances aligned.
|
| 67 |
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
| FluencyBank (stuttered, 5,487 words) | 77.1 | 25.0 | 0.743 | 58.0 | 82.4 |
|
| 73 |
|
| 74 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
|
| 76 |
-
|
| 77 |
-
|
|
| 78 |
-
|
| 79 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 80 |
|
| 81 |
-
|
| 82 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
|
| 84 |
`nyra_forced_aligner_en_pro` delivers roughly the same accuracy on clean speech
|
| 85 |
-
and is meaningfully more robust to noise and adverse acoustic conditions
|
| 86 |
-
not publicly released and available on request.
|
| 87 |
|
| 88 |
-
|
|
|
|
| 89 |
|
| 90 |
-
|
| 91 |
-
- 10 ms frame grid; median boundary error is ~14 ms on conversational speech.
|
| 92 |
-
- Like other GMM-HMM aligners, word boundaries adjacent to long pauses tend to
|
| 93 |
-
extend ~25 ms into the pause.
|
| 94 |
-
- The transcript must be verbatim for best results; missing or extra words are
|
| 95 |
-
absorbed into neighbouring words.
|
| 96 |
|
| 97 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 98 |
|
| 99 |
-
|
| 100 |
-
|---|---|
|
| 101 |
-
| `model.npz` | GMM parameters, transition model, triphone->pdf table, LDA/MLLT matrices |
|
| 102 |
-
| `meta.json` | topology, phone table, feature chain, decoding config |
|
| 103 |
-
| `lexicon.txt` | 41k-word pronunciation lexicon |
|
| 104 |
-
| `phone_map.json` | espeak IPA -> phone mapping used for out-of-lexicon words |
|
| 105 |
-
| `wavlm/` | WavLM-large weights (Microsoft; redistributed from `microsoft/wavlm-large`) |
|
| 106 |
|
| 107 |
-
|
| 108 |
|
| 109 |
-
The
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
improve models intended for commercial use — requires a commercial license**
|
| 115 |
-
from nyra health GmbH (licensing@nyra-labs.com).
|
| 116 |
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 120 |
|
| 121 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 122 |
|
| 123 |
-
|
|
|
|
|
|
| 3 |
license_name: nyra-health-non-commercial-research-license
|
| 4 |
license_link: LICENSE.md
|
| 5 |
language: [en]
|
| 6 |
+
library_name: nyra-forced-aligner
|
| 7 |
base_model: microsoft/wavlm-large
|
| 8 |
tags: [forced-alignment, speech, phonetics, wavlm, verbatim, disfluency, timestamps]
|
| 9 |
---
|
| 10 |
|
| 11 |
+
# nyra-forced-aligner
|
| 12 |
|
| 13 |
+
[](https://huggingface.co/nyralabs)
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
+
**The most accurate word timing you can get for conversational speech:
|
| 16 |
+
verbatim-aware, noise-robust, and fast.**
|
| 17 |
|
| 18 |
+
[Blog post](https://nyra-labs.com/research/nyra-forced-aligner) ·
|
| 19 |
+
[Paper](#) <!-- TODO: ICASSP 2026 link --> ·
|
| 20 |
+
[Models](https://huggingface.co/nyralabs) ·
|
| 21 |
+
[CrisperWhisper](https://github.com/nyrahealth/CrisperWhisper)
|
| 22 |
|
| 23 |
+
A forced aligner answers one question: given a recording and its transcript,
|
| 24 |
+
*when* was each word said? Most aligners were built for read speech and clean
|
| 25 |
+
audio, and they quietly fall apart on the material people actually need to
|
| 26 |
+
align: conversations full of `um`s, laughter, restarts and half-finished words,
|
| 27 |
+
recorded in rooms that are not studios.
|
| 28 |
+
|
| 29 |
+
`nyra-forced-aligner` was built for exactly that material.
|
| 30 |
+
|
| 31 |
+
- **Vocal sounds are alignable.** Fillers (`[UH]`, `[UM]`), laughter, breaths,
|
| 32 |
+
coughs and word cut-offs (`th-`) are units of the model like any phone,
|
| 33 |
+
so they get timestamps of their own instead of being smeared into the
|
| 34 |
+
neighbouring words.
|
| 35 |
+
- **Made for CrisperWhisper transcripts.** Trained on the verbatim
|
| 36 |
+
transcription convention of
|
| 37 |
+
[CrisperWhisper 2.0](https://github.com/nyrahealth/CrisperWhisper): feed
|
| 38 |
+
its output straight in and every filler, repetition and event token is
|
| 39 |
+
aligned.
|
| 40 |
+
- **Accurate where it matters.** 19.9 ms mean word-boundary error on
|
| 41 |
+
conversational English (Buckeye): 21% lower than the best configuration of
|
| 42 |
+
the Montreal Forced Aligner, half the error of the best ASR-with-timestamps
|
| 43 |
+
system, and the only system whose error tail drops below 1% of words before
|
| 44 |
+
100 ms.
|
| 45 |
+
- **Noise-robust.** Every utterance stays aligned down to 0 dB SNR, where
|
| 46 |
+
classical aligners lose or mangle most of them. A Pro model goes further for
|
| 47 |
+
difficult recordings.
|
| 48 |
+
- **Fast and longform.** Batched WavLM features, GPU emissions and a compiled
|
| 49 |
+
beam Viterbi: eight minutes of audio align in about three seconds on one
|
| 50 |
+
GPU.
|
| 51 |
+
|
| 52 |
+
## Install
|
| 53 |
+
|
| 54 |
+
```bash
|
| 55 |
+
pip install nyra-forced-aligner
|
| 56 |
+
|
| 57 |
+
# espeak-ng is needed to pronounce words outside the built-in 41k-word lexicon:
|
| 58 |
+
# apt install espeak-ng (macOS: brew install espeak-ng)
|
| 59 |
+
```
|
| 60 |
+
|
| 61 |
+
Runs on CPU, but a CUDA GPU is strongly recommended (WavLM-large front-end).
|
| 62 |
+
The model is downloaded from the Hugging Face Hub on first use.
|
| 63 |
+
|
| 64 |
+
## Quickstart
|
| 65 |
|
| 66 |
```python
|
| 67 |
from nyra_align import Aligner
|
| 68 |
|
| 69 |
+
aligner = Aligner() # nyralabs/nyra_forced_aligner_en
|
| 70 |
+
# aligner = Aligner(pro=True) # noise-robust model
|
| 71 |
+
|
| 72 |
+
result = aligner.align("interview.wav",
|
| 73 |
+
"so i [UM] i went there on on thursday [laughter]")
|
| 74 |
|
| 75 |
for w in result.words:
|
| 76 |
+
print(f"{w.start:7.3f} {w.end:7.3f} {w.word:<12s} {w.type}") # w / f(iller) / s(ound) / c(ut-off)
|
| 77 |
+
|
| 78 |
+
result.to_textgrid("interview.TextGrid") # Praat
|
| 79 |
result.to_json("interview.json")
|
| 80 |
```
|
| 81 |
|
| 82 |
+
Any audio length works. For the best timings follow the
|
| 83 |
+
[transcript conventions](#transcript-conventions) below.
|
| 84 |
|
| 85 |
+
Command line, single file or a whole corpus of paired `.wav` + `.txt`/`.lab`
|
| 86 |
+
files (MFA-style layout):
|
| 87 |
|
| 88 |
```bash
|
| 89 |
+
nyra-align interview.wav --text-file interview.txt --out interview.TextGrid
|
| 90 |
+
nyra-align corpus_dir/ --pro --out-dir aligned/ --format textgrid
|
| 91 |
+
```
|
| 92 |
+
|
| 93 |
+
### Transcript conventions
|
| 94 |
+
|
| 95 |
+
The aligner can only place what the transcript contains, so for the best
|
| 96 |
+
timings the transcript should be **verbatim: write what you hear.**
|
| 97 |
+
|
| 98 |
+
- **Cut-offs.** Interrupted words and word fragments are marked with a
|
| 99 |
+
trailing hyphen: `th-`, `w-`, `resched-`.
|
| 100 |
+
- **Fillers.** Filled pauses are bracketed and upper-cased: `[UH]`, `[UM]`.
|
| 101 |
+
- **Vocal sound events.** Non-speech sounds are bracketed: `[laughter]`,
|
| 102 |
+
`[cough]`, `[breath]`, `[sigh]`, `[sniff]`, `[lipsmack]`,
|
| 103 |
+
`[throatclearing]`, `[yawn]`, `[noise]`.
|
| 104 |
+
- **Numbers, dates, times and emails** are written out exactly as spoken
|
| 105 |
+
(`March third at nine thirty`), not as digits or symbols.
|
| 106 |
+
- **Everything else** is transcribed word for word. Normal casing and
|
| 107 |
+
punctuation are fine (the aligner ignores both); repetitions, false starts,
|
| 108 |
+
fillers, fragments and sounds are all left in place.
|
| 109 |
+
|
| 110 |
+
```text
|
| 111 |
+
so we we need to, to reschedule the th- Thursday meeting to [UH] March third at nine thirty [laughter]
|
| 112 |
```
|
| 113 |
|
| 114 |
+
What happens when a transcript deviates: a word that was spoken but is missing
|
| 115 |
+
from the transcript gets absorbed into its neighbours, and a word in the
|
| 116 |
+
transcript that was never spoken is squeezed into a few frames, so both shift
|
| 117 |
+
the surrounding boundaries. Event tags are matched in any casing; a tag the
|
| 118 |
+
model does not know (`[music]`) is left out and listed in `result.skipped`,
|
| 119 |
+
as is any token that cannot be pronounced.
|
| 120 |
|
| 121 |
+
### With CrisperWhisper 2.0
|
| 122 |
|
| 123 |
+
[CrisperWhisper 2.0](https://github.com/nyrahealth/CrisperWhisper) produces
|
| 124 |
+
transcripts in exactly this convention, so its verbatim output can be aligned
|
| 125 |
+
as is:
|
|
|
|
|
|
|
| 126 |
|
| 127 |
+
```python
|
| 128 |
+
from crisperwhisper import CrisperWhisperModel
|
| 129 |
+
from nyra_align import Aligner
|
| 130 |
+
|
| 131 |
+
audio = "interview.wav"
|
| 132 |
+
transcript = CrisperWhisperModel("large").transcribe(audio, language="en").text
|
| 133 |
+
result = Aligner().align(audio, transcript)
|
| 134 |
+
|
| 135 |
+
for w in result.words:
|
| 136 |
+
print(f"{w.start:7.3f} {w.end:7.3f} {w.word}")
|
| 137 |
+
result.to_textgrid("interview.TextGrid")
|
| 138 |
+
```
|
| 139 |
+
|
| 140 |
+
### Models
|
| 141 |
+
|
| 142 |
+
| Shorthand | Hugging Face ID | Use it for |
|
| 143 |
+
|-----------|-----------------|------------|
|
| 144 |
+
| `Aligner()` (default) | `nyralabs/nyra_forced_aligner_en` | Clean and lightly noisy audio; most accurate on clean speech |
|
| 145 |
+
| `Aligner(pro=True)` | `nyralabs/nyra_forced_aligner_en_pro` | **Pro**: roughly the same accuracy on clean speech, meaningfully more robust to noise and adverse acoustic conditions. Not publicly released; available on request |
|
| 146 |
|
| 147 |
+
Only English is available at the moment; `Aligner(language="de")` raises
|
| 148 |
+
`NotImplementedError`.
|
| 149 |
|
| 150 |
+
## Performance
|
|
|
|
| 151 |
|
| 152 |
+
Mean absolute word-boundary error on **Buckeye** (conversational English,
|
| 153 |
+
2,008 test recordings, 80,475 words), lower is better. Every system is scored
|
| 154 |
+
with the same protocol: DP word matching on canonicalized forms, laughter
|
| 155 |
+
tokens excluded, MAE = mean over words of (|onset error| + |offset error|)/2.
|
|
|
|
| 156 |
|
| 157 |
+
| # | System | Type | MAE (ms) | MedAE | ≤25 ms | ≤50 ms | ≤100 ms | mIoU |
|
| 158 |
+
|--:|--------|------|---------:|------:|-------:|-------:|--------:|-----:|
|
| 159 |
+
| 1 | **nyra_forced_aligner_en** | forced aligner | **19.9** | 13.9 | **79.5%** | **93.7%** | **98.3%** | **0.830** |
|
| 160 |
+
| 2 | Montreal Forced Aligner `english_us_arpa`, wide beam¹ | forced aligner | 25.3 | 13.0 | 76.7% | 92.2% | 97.3% | 0.826 |
|
| 161 |
+
| 3 | Montreal Forced Aligner `english_mfa`, wide beam¹ | forced aligner | 35.6 | 15.0 | 71.3% | 87.5% | 95.2% | 0.790 |
|
| 162 |
+
| 4 | CrisperWhisper 2.0 | ASR | 38.7 | 25.0 | 54.2% | 78.7% | 91.7% | 0.717 |
|
| 163 |
+
| 5 | Qwen3-ForcedAligner-0.6B | forced aligner | 39.5 | 25.0 | 52.8% | 85.9% | 96.6% | 0.711 |
|
| 164 |
+
| 6 | xAI Grok Speech-to-Text | ASR | 45.5 | 34.5 | 58.3% | 82.8% | 94.0% | 0.630 |
|
| 165 |
+
| 7 | Montreal Forced Aligner `english_us_arpa`, default beam | forced aligner | 48.0 | 13.5 | 75.2% | 90.4% | 95.6% | 0.810 |
|
| 166 |
+
| 8 | Montreal Forced Aligner `english_mfa`, default beam | forced aligner | 48.9 | 15.5 | 70.7% | 86.7% | 94.3% | 0.782 |
|
| 167 |
+
| 9 | MMS-FA (torchaudio) | forced aligner | 53.0 | 34.9 | 57.5% | 81.3% | 92.8% | 0.620 |
|
| 168 |
+
| 10 | ElevenLabs Scribe v2 | ASR | 57.8 | 50.0 | 47.6% | 75.8% | 94.8% | 0.540 |
|
| 169 |
+
| 11 | Charsiu | forced aligner | 70.3 | 28.0 | 51.8% | 68.5% | 85.1% | 0.649 |
|
| 170 |
+
| 12 | NeMo Forced Aligner | forced aligner | 72.6 | 62.5 | 26.8% | 49.8% | 79.3% | 0.443 |
|
| 171 |
+
| 13 | WhisperX (wav2vec2 alignment) | forced aligner | 80.9 | 61.5 | 35.1% | 68.1% | 94.1% | 0.456 |
|
| 172 |
+
| 14 | CTC-Segmentation | forced aligner | 81.0 | 36.0 | 41.5% | 69.5% | 89.3% | 0.635 |
|
| 173 |
+
| 15 | Deepgram Nova-3 | ASR | 86.9 | 71.0 | 18.1% | 36.4% | 68.9% | 0.496 |
|
| 174 |
+
| 16 | Cartesia Ink-Whisper | ASR | 124.8 | 46.3 | 36.0% | 60.4% | 82.8% | 0.598 |
|
| 175 |
|
| 176 |
+
<sub>**MAE** and **MedAE** are over the per-word error (|onset error| +
|
| 177 |
+
|offset error|)/2. The **≤ t ms** columns follow the convention of Rousso et
|
| 178 |
+
al. (2024): the share of words whose *end* timestamp is within t of the
|
| 179 |
+
reference (one boundary, not the per-word mean). **Forced aligners** are given
|
| 180 |
+
the reference words and only predict their timing; every word is scored. **ASR** systems produce their own words, which
|
| 181 |
+
are text-aligned to the reference, and only matched words are scored
|
| 182 |
+
(88–93% of words), so their timing is judged independently of their
|
| 183 |
+
transcription errors. ¹ MFA's shipped default beam (10) prunes the correct
|
| 184 |
+
path on conversational speech; the "wide beam" rows use `beam 400 / retry
|
| 185 |
+
4000`, the best configuration we found for it.</sub>
|
| 186 |
|
| 187 |
+
The mean hides where the difference really is. Averaged over all words, most
|
| 188 |
+
systems are within a few tens of milliseconds; the gap is in the **tail**, the
|
| 189 |
+
words that are misaligned by 100 ms or more, which is what breaks downstream
|
| 190 |
+
use. Each curve shows the share of words whose per-word error (mean of the
|
| 191 |
+
onset and offset errors, the quantity behind MAE) is larger than *t*; note
|
| 192 |
+
this is not the single-boundary ≤ t ms column of the table:
|
| 193 |
+
|
| 194 |
+

|
| 195 |
+
|
| 196 |
+
### Under noise
|
| 197 |
+
|
| 198 |
+
The same Buckeye recordings with added noise, at signal-to-noise ratios from
|
| 199 |
+
20 dB down to 0 dB (MAE in ms). MUSAN is real noise recordings; Gaussian is
|
| 200 |
+
white noise, which the model never saw in training.
|
| 201 |
+
|
| 202 |
+
| System | MUSAN 20 | 10 | 5 | 0 dB | Gauss 20 | 10 | 5 | 0 dB |
|
| 203 |
+
|--------|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 204 |
+
| **nyra_forced_aligner_en** | 20.0 | 22.2 | 26.9 | 56.3 | 20.2 | 25.3 | 38.9 | 110.9 |
|
| 205 |
+
| CrisperWhisper 2.0² | 38.7 | 42.4 | 48.0 | 55.7 | 40.1 | 50.9 | 67.3 | 103.4 |
|
| 206 |
+
| Qwen3-ForcedAligner-0.6B | 40.6 | 54.3 | 78.7 | 153.9 | 43.4 | 85.2 | 153.9 | 339.2 |
|
| 207 |
+
| MMS-FA (torchaudio) | 52.8 | 54.9 | 57.7 | 67.0 | 52.0 | 55.2 | 61.0 | 86.1 |
|
| 208 |
+
| Montreal Forced Aligner `english_us_arpa`, wide beam | 33.3 | 65.3 | 127.2 | 281.2 | 36.2 | 74.4 | 156.3 | 416.9 |
|
| 209 |
+
| WhisperX (wav2vec2 alignment) | 81.7 | 90.4 | 138.1 | 468.5 | 84.2 | 126.3 | 415.2 | 2720.6 |
|
| 210 |
+
|
| 211 |
+
<sub>² CrisperWhisper 2.0 is an ASR system: at 0 dB it still recognises only
|
| 212 |
+
29% of the words, and only those are scored. All forced aligners above keep
|
| 213 |
+
100% of words at every SNR.</sub>
|
| 214 |
|
| 215 |
`nyra_forced_aligner_en_pro` delivers roughly the same accuracy on clean speech
|
| 216 |
+
and is meaningfully more robust to noise and adverse acoustic conditions.
|
|
|
|
| 217 |
|
| 218 |
+
Results on read speech (TIMIT) and stuttered speech (FluencyBank), the full
|
| 219 |
+
set of ablations, and how the models are trained are in the paper.
|
| 220 |
|
| 221 |
+
## What else is in the box
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 222 |
|
| 223 |
+
| Option | What it does |
|
| 224 |
+
|--------|--------------|
|
| 225 |
+
| `Aligner(pro=True)` | Noise-robust model |
|
| 226 |
+
| `Aligner(device="cuda" / "cpu")` | Device selection (auto by default) |
|
| 227 |
+
| `Aligner(precision="fp32")` | WavLM in fp32 instead of fp16; bit-exact parity with the reference implementation (fp16 moves ≤30 ms on ~0.5% of word boundaries) |
|
| 228 |
+
| `Aligner(wavlm_batch=16)` | Batch size for the 30 s WavLM windows on long audio (tune to VRAM) |
|
| 229 |
+
| `aligner.align_features(feats, text)` | Re-align a transcript against cached features (decode is ~6 ms per 3.5 s utterance) |
|
| 230 |
+
| `result.events` | Just the filler / laughter / cut-off tokens |
|
| 231 |
+
| `result.silences` | Silence segments between words |
|
| 232 |
|
| 233 |
+
## How it works
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 234 |
|
| 235 |
+
The full story is in the deep-dive post:
|
| 236 |
|
| 237 |
+
- [The nyra forced aligner: precise word timing for real conversational speech](https://nyra-labs.com/research/nyra-forced-aligner).
|
| 238 |
+
Why classical aligners break on spontaneous and noisy speech, how frozen
|
| 239 |
+
WavLM features, vocal-event units and a noise-pooled projection fix it, and
|
| 240 |
+
how the model bootstraps itself from verbatim transcripts without any
|
| 241 |
+
manual boundary labels.
|
|
|
|
|
|
|
| 242 |
|
| 243 |
+
In short: WavLM-large hidden states are projected by a supervised LDA onto 40
|
| 244 |
+
dimensions and modelled by a triphone GMM-HMM whose inventory holds 15
|
| 245 |
+
vocal-event units next to the 39 English phones. At inference the package
|
| 246 |
+
builds the alignment graph for the transcript (optional silence between words,
|
| 247 |
+
full context across word boundaries) and runs a compiled beam Viterbi over the
|
| 248 |
+
model's likelihoods, with a beam ladder so that no utterance is ever silently
|
| 249 |
+
dropped. Words outside the built-in lexicon are phonemized with espeak through
|
| 250 |
+
the same mapping used at training time.
|
| 251 |
+
|
| 252 |
+
## License
|
| 253 |
|
| 254 |
+
The model files are released under the
|
| 255 |
+
[nyra health Non-Commercial Research License](https://huggingface.co/nyralabs/nyra_forced_aligner_en/blob/main/LICENSE.md):
|
| 256 |
+
free for research and other non-commercial use. This license covers the
|
| 257 |
+
model **and the alignments it produces**: any commercial use of the
|
| 258 |
+
timestamps, and any use of them to train, fine-tune or otherwise improve
|
| 259 |
+
models intended for commercial use, requires a commercial license. The Pro
|
| 260 |
+
model is available under commercial license only. For commercial licensing
|
| 261 |
+
of either, [contact nyra health](mailto:licensing@nyra-labs.com).
|
| 262 |
|
| 263 |
+
The model repositories bundle Microsoft's WavLM-large weights, redistributed
|
| 264 |
+
unchanged under their upstream license.
|
docs/buckeye_ccdf.png
ADDED
|
Git LFS Details
|