owensong commited on
Commit
b4edf7f
·
verified ·
1 Parent(s): a129c4a

Clarify benchmark documentation

Browse files
Files changed (2) hide show
  1. docs/EVALUATION.md +10 -11
  2. release_manifest.json +2 -2
docs/EVALUATION.md CHANGED
@@ -7,10 +7,9 @@ cards show the headline results; this document defines the datasets, systems,
7
  normalization, uncertainty, exclusions, runtime boundaries, and raw artifacts
8
  behind those results.
9
 
10
- > **Interpret results by axis, not as one universal score.** Intelligibility,
11
- > predicted naturalness, human preference, model footprint, and runtime measure
12
- > different properties. Inflect does not combine them into a proprietary
13
- > aggregate.
14
 
15
  ## Evidence at a glance
16
 
@@ -25,13 +24,13 @@ behind those results.
25
 
26
  ![Two-ASR semantic WER](../assets/evidence/asr-consensus.svg)
27
 
28
- ## Comparator policy
29
-
30
- The comparison set intentionally uses serious compact or on-device systems:
31
- KittenTTS Nano, Piper Low, and Supertonic 3. It does not use deliberately weak
32
- baselines. Every comparator in this package-level study has a larger deployable
33
- weight footprint than both Inflect releases. That makes the comparison
34
- demanding, but it does not mean one benchmark proves universal superiority.
35
 
36
  Voice variants are separate rows when voice identity can affect quality or
37
  intelligibility. Kitten and Piper may be shown as equal-weight two-voice means
 
7
  normalization, uncertainty, exclusions, runtime boundaries, and raw artifacts
8
  behind those results.
9
 
10
+ Read each metric separately: WER measures intelligibility, UTMOS22 predicts
11
+ naturalness, listening tests record preference, and runtime measures deployment
12
+ cost. None of them is an overall quality score.
 
13
 
14
  ## Evidence at a glance
15
 
 
24
 
25
  ![Two-ASR semantic WER](../assets/evidence/asr-consensus.svg)
26
 
27
+ ## Compared systems
28
+
29
+ The benchmark includes KittenTTS Nano, Piper Low, and Supertonic 3 because they
30
+ are established compact or local TTS systems. In this package-level comparison,
31
+ each has a larger deployable weight footprint than both Inflect releases. That
32
+ provides useful size context, but no single result establishes overall
33
+ superiority.
34
 
35
  Voice variants are separate rows when voice identity can affect quality or
36
  intelligibility. Kitten and Piper may be shown as equal-weight two-voice means
release_manifest.json CHANGED
@@ -84,8 +84,8 @@
84
  },
85
  {
86
  "path": "docs/EVALUATION.md",
87
- "bytes": 12599,
88
- "sha256": "b350f2990dc6a192211bd065f5f28920cb513507981f42d76e0ad0bc1e6e724b"
89
  },
90
  {
91
  "path": "docs/EXPORTS.md",
 
84
  },
85
  {
86
  "path": "docs/EVALUATION.md",
87
+ "bytes": 12488,
88
+ "sha256": "9e0d4c1811ca9959529db4e7b6907ca931f094dfd910dbe7f61b550fbc732f17"
89
  },
90
  {
91
  "path": "docs/EXPORTS.md",