# Final evaluation protocol ## Matched corpus All automated headline results use 500 identical unseen English prompts per system. The set combines 200 manually audited stress prompts with 300 deterministic CMU Arctic sentences. Exact normalized text was excluded against 87,362 training transcripts. The frozen prompt manifest and exclusion report are included under `evaluation/final/raw/`. ## Systems - Inflect-Micro-v2 and Inflect-Nano-v2, fixed release checkpoints. - KittenTTS Nano, Bruno and Hugo reported separately. - Piper low, Ryan and Danny reported separately. - Supertonic 3 James at 3 and 8 steps, reported separately. Family averages for Kitten and Piper are equal-weight macro averages of the two named voices. The two Supertonic step counts are never pooled. ## Intelligibility Headline semantic WER uses Whisper-large-v3. A second complete pass with wav2vec2 Large LV-60K tests whether conclusions are sensitive to one ASR family. References and hypotheses are normalized into equivalent spoken English forms so formatting differences such as `12.6` versus `twelve point six` are not automatically counted as speech errors. Both reports include per-clip rows, prompt-category breakdowns, and 10,000-sample bootstrap intervals. ASR-family results are never averaged into a single headline value. ## Predicted quality UTMOS22 is run over every generated clip with 10,000-sample bootstrap intervals. It is a learned quality predictor, not human MOS, and is reported separately from WER and human preference. ## Runtime Runtime reports use one named host, the same 50-prompt subset, warm end-to-end synthesis, one isolated process per system, each public runtime's default thread/provider behavior, and separate model-load timing. RTF, audio-seconds per wall-second, median latency, p95 latency, and per-prompt rows are retained. Cross-machine numbers are not mixed. ## Human evidence The community study was blind, randomized left/right order, and stored pairwise preferences rather than user identity. Preference rate is `(wins + 0.5 × ties) / appearances`. It is descriptive community evidence, not a formal laboratory MOS study. ## Reporting rules No aggregate score combines WER, UTMOS22, speed, size, and human preference. Confidence intervals, raw reports, model weight footprint, and known limitations accompany headline values.