Inflect-Nano-v2 is the portability-focused member of Inflect v2. The complete deployable system has **3,966,721 parameters**; there is no second vocoder download and no hidden large language model at inference time.
## Listen First
Every clip below uses held-out text rather than a training-audio reconstruction.
| Prompt type | Audio |
| --- | --- |
| Conversational | |
| Punctuation | |
| Numbers | |
| Names and places | |
| Technical | |
## At a Glance
| | Inflect-Nano-v2 |
| --- | ---: |
| Complete deployable parameters | **3,966,721** |
| FP32 model weights | **15.97 MB** |
| Sample rate | **24 kHz** |
| Voice | **1 fixed English male voice** |
| External vocoder | **None** |
| Release format | **PyTorch FP32** |
| Long text | **Punctuation-aware chunking** |
Inflect v2 is released in two deliberately different sizes. **Micro** spends its larger budget on waveform capacity and hidden width; **Nano** prioritizes footprint while preserving the same public runtime and text frontend. Both numbers above count the complete deployable generator, including the waveform decoder.
## Human Listening

Inflect-Nano-v2 recorded a **63.9% preference rate** (22 wins, 12 losses, 2 ties) in the final anonymous community study. The full study contains **147 pairwise decisions** across the published lineup. Ties count as half a win. This is descriptive community evidence, not a formal MOS study; pair coverage and vote counts differ by system.
## Evaluation
### Quality versus footprint

### Intelligibility

### Cross-ASR robustness

### CPU throughput

### Quality and speed

### Error analysis

### Complete model footprint

| System | Weights | UTMOS22 ↑ | Semantic WER ↓ |
| --- | ---: | ---: | ---: |
| **Inflect-Micro-v2** | 37.5 MB | 4.395 | 1.93% |
| **Inflect-Nano-v2** | 16.0 MB | 4.386 | 2.48% |
| KittenTTS Nano · Bruno | 56.8 MB | 4.217 | 0.97% |
| KittenTTS Nano · Hugo | 56.8 MB | 4.191 | 1.15% |
| Piper · Ryan Low | 63.1 MB | 4.289 | 2.33% |
| Piper · Danny Low | 63.1 MB | 4.195 | 1.81% |
| Supertonic 3 · James · 3-step | 398.1 MB | 2.471 | 692.82% |
| Supertonic 3 · James · 8-step | 398.1 MB | 4.295 | 622.52% |
On the 500-prompt matched set, Inflect-Nano-v2 measured **4.386 UTMOS22** (95% bootstrap CI 4.372–4.399) and **2.48% semantic WER** (95% CI 1.95–3.07). UTMOS22 is an automated quality predictor, not human MOS. WER measures transcript recoverability, not naturalness.
Under the prespecified Whisper-large-v3 protocol, Supertonic produced insertion-heavy hypotheses and landed off-scale at **692.82%** (3-step) and **622.52%** (8-step). The independent wav2vec2 pass scored the same clips at **13.47%** and **3.76%**, respectively. This large disagreement is reported as ASR-family sensitivity, not collapsed into an aggregate or interpreted as a direct measure of Supertonic intelligibility. Exact hypotheses and per-clip errors ship in the raw reports.
Protocol and raw evidence
- 500 identical unseen English prompts per system; 4,000 generated clips total.
- The set combines a 200-prompt stress corpus with 300 deterministic CMU Arctic phonetic-balance sentences.
- Exact-text exclusion was run against 87,362 training transcripts.
- Headline semantic WER uses Whisper-large-v3 and the disclosed English normalizer; wav2vec2 LV-60K is included as an independent ASR-sensitivity check.
- UTMOS22 uses `tarepan/SpeechMOS` v1.2.0.
- Headline intervals use 10,000 bootstrap samples.
- Kitten and Piper voices are reported separately; equal-weight family averages are in `evaluation/final/evidence-summary.json`.
- Supertonic 3-step and 8-step are never pooled.
- Per-clip outputs, category breakdowns, runtime rows, and signal diagnostics are under `evaluation/final/`.
## Quickstart
```bash
pip install -r requirements.txt
python inference.py \
--model-dir . \
--device cpu \
--text "A small voice can still have something meaningful to say." \
--output sample.wav
```
```python
import sys
from huggingface_hub import snapshot_download
model_dir = snapshot_download("owensong/Inflect-Nano-v2")
sys.path.insert(0, model_dir)
from inference import InflectTTS
tts = InflectTTS(model_dir, device="cpu")
tts.save(
"The complete speech model runs locally.",
"sample.wav",
speed=1.0,
variation=0.667,
seed=7,
)
```
The packaged runtime accepts `speed` from `0.5` to `2.0`, `variation` from `0.0` to `1.0`, and a deterministic integer seed. Long input is split at punctuation-aware boundaries and concatenated with controlled pauses.
### Controls
| Control | Default | Effect |
| --- | ---: | --- |
| `speed` | `1.0` | Speech rate; supported range `0.5`–`2.0` |
| `variation` | `0.667` | Stochastic latent variation; lower is steadier, higher is more variable |
| `seed` | `0` | Reproducible sampling seed; incremented per long-text chunk |
Use a fixed seed when comparing systems. A different seed is a different stochastic sample, not a different checkpoint.
## Architecture
Inflect v2 is a **parameter-efficient VITS-family, end-to-end text-to-waveform speech generator** with a phoneme frontend, monotonic alignment, stochastic latent synthesis, and an integrated neural waveform decoder.
- punctuation-preserving English normalization and stress-marked phonemes
- compact Transformer text encoder and duration model
- latent prior and residual coupling flow
- integrated adversarial waveform decoder at 24 kHz
- inference-only checkpoint with training discriminators and optimizer state removed
| Component | Value |
| --- | ---: |
| Latent channels | 128 |
| Text hidden channels | 72 |
| Text encoder layers / heads | 3 / 2 |
| Feed-forward channels | 384 |
| Flow coupling blocks | 4 |
| Initial decoder channels | 192 |
| Upsample rates | 8, 8, 2, 2 |
| Training segment | 16,384 samples |
The model card describes the deployable architecture. The private corpus-construction and optimization recipe is not part of this open-weight release.
## Runtime
Warm end-to-end CPU inference on **AMD EPYC 7K62 48-Core Processor**, with **192 logical CPUs visible** to the isolated process and runtime-default thread settings:
| RTF ↓ | Audio / wall time ↑ | Median latency | p95 latency | Cold load |
| ---: | ---: | ---: | ---: | ---: |
| 0.5400 | 1.85× | 2.199 s | 3.850 s | 0.82 s |
Runtime varies with processor, PyTorch build, thread policy, text length, and operating system. Use the raw report rather than extrapolating these measurements to untested devices.
## Fine-tuning
The initial release is inference-first. The checkpoint is structurally compatible with voice and language adaptation, but the private training recipe and corpus-generation pipeline are not included. A tested public adaptation toolkit should be treated as a separate release rather than inferred from the inference package. See `docs/FINETUNING.md` for the supported status and data contract.
## Repository Layout
| Path | Purpose |
| --- | --- |
| `model.pth` | Inference-only generator checkpoint |
| `config.json` | Deployable architecture and audio configuration |
| `inference.py` | Public Python API and CLI |
| `inflect_vits_frontend.py` | Normalization, phonemization, and punctuation frontend |
| `runtime/` | Self-contained model implementation required for inference |
| `samples/` | Held-out example generations |
| `evaluation/final/` | Final matched benchmark reports and protocol artifacts |
| `release_manifest.json` | File sizes and SHA-256 integrity hashes |
## Limitations
- English only and one fixed male voice. This is not a zero-shot voice-cloning model.
- Unfamiliar phrasing can become flatter or less expressive.
- Numbers, abbreviations, homographs, and uncommon names depend strongly on frontend normalization and context.
- Long passages are synthesized in chunks, so transitions can differ from a single long-form model pass.
- Stochastic variation can change pronunciation and timing; use a fixed seed for reproducibility.
- Learned MOS predictors can reward or penalize spectral traits differently from human listeners.
- Not validated for medical, legal, emergency, or accessibility-critical communication.
## Responsible Use
Do not use the included voice to impersonate a real person, deceive listeners, or create fraudulent content. Disclose synthetic speech where the context could otherwise mislead. Users are responsible for applicable laws and the license.
## Integrity and License
`release_manifest.json` records file sizes and SHA-256 hashes. Original Inflect code and weights in this repository are licensed under Apache-2.0; bundled third-party components retain their own notices in `THIRD_PARTY_NOTICES.md`.
## Citation
```bibtex
@software{song2026inflectnanov2,
author = {Owen Song},
title = {Inflect-Nano-v2: Complete Local Text-to-Waveform TTS Under 4M Parameters},
year = {2026},
url = {https://huggingface.co/owensong/Inflect-Nano-v2}
}
```