Deployment notes
Installation
Use Python 3.10 or newer in a clean virtual environment, install requirements.txt, then run inference.py from the model repository. The package is tested with PyTorch CPU and CUDA execution; device-specific performance depends on the local PyTorch build.
Reproducibility
Keep the checkpoint, config.json, frontend files, and runtime/ directory together. Record Inflect-Nano-v2, the release revision, speed, variation, and seed with generated artifacts. Verify files against release_manifest.json when moving the package between machines.
Long text
The runtime splits long text at sentence punctuation and then at safe intra-sentence boundaries. Chunks are synthesized independently and joined with punctuation-dependent silence. This prevents unbounded inference but means a long paragraph is not one globally planned prosody pass.
Production boundaries
- Instantiate
InflectTTSonce and reuse it; loading per request wastes time. - Serialize calls to one model instance unless the surrounding application explicitly manages device memory and random seeds.
- Validate empty or untrusted input before synthesis.
- Set a fixed seed for deterministic caching and tests.
- Measure on the actual target device before claiming real-time operation.
Troubleshooting
- Missing phonemizer errors: reinstall
phonemizerandespeakng-loaderfromrequirements.txt. - Unexpected pronunciation: spell out ambiguous abbreviations or add context; the release frontend is English-only.
- Flat or unstable delivery: keep
variationnear the default before changing speed. - Long pauses: inspect source punctuation because chunk boundaries intentionally follow punctuation.
- Slow first request: model loading is measured separately from warm synthesis and should not happen per utterance.