# Deployment notes ## Installation Use Python 3.10 or newer in a clean virtual environment, install `requirements.txt`, then run `inference.py` from the model repository. The package is tested with PyTorch CPU and CUDA execution; device-specific performance depends on the local PyTorch build. ## Reproducibility Keep the checkpoint, `config.json`, frontend files, and `runtime/` directory together. Record `Inflect-Nano-v2`, the release revision, `speed`, `variation`, and `seed` with generated artifacts. Verify files against `release_manifest.json` when moving the package between machines. ## Long text The runtime splits long text at sentence punctuation and then at safe intra-sentence boundaries. Chunks are synthesized independently and joined with punctuation-dependent silence. This prevents unbounded inference but means a long paragraph is not one globally planned prosody pass. ## Production boundaries - Instantiate `InflectTTS` once and reuse it; loading per request wastes time. - Serialize calls to one model instance unless the surrounding application explicitly manages device memory and random seeds. - Validate empty or untrusted input before synthesis. - Set a fixed seed for deterministic caching and tests. - Measure on the actual target device before claiming real-time operation. ## Troubleshooting - Missing phonemizer errors: reinstall `phonemizer` and `espeakng-loader` from `requirements.txt`. - Unexpected pronunciation: spell out ambiguous abbreviations or add context; the release frontend is English-only. - Flat or unstable delivery: keep `variation` near the default before changing speed. - Long pauses: inspect source punctuation because chunk boundaries intentionally follow punctuation. - Slow first request: model loading is measured separately from warm synthesis and should not happen per utterance.