File size: 1,882 Bytes
125985b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
# Deployment notes

## Installation

Use Python 3.10 or newer in a clean virtual environment, install `requirements.txt`, then run `inference.py` from the model repository. The package is tested with PyTorch CPU and CUDA execution; device-specific performance depends on the local PyTorch build.

## Reproducibility

Keep the checkpoint, `config.json`, frontend files, and `runtime/` directory together. Record `Inflect-Nano-v2`, the release revision, `speed`, `variation`, and `seed` with generated artifacts. Verify files against `release_manifest.json` when moving the package between machines.

## Long text

The runtime splits long text at sentence punctuation and then at safe intra-sentence boundaries. Chunks are synthesized independently and joined with punctuation-dependent silence. This prevents unbounded inference but means a long paragraph is not one globally planned prosody pass.

## Production boundaries

- Instantiate `InflectTTS` once and reuse it; loading per request wastes time.
- Serialize calls to one model instance unless the surrounding application explicitly manages device memory and random seeds.
- Validate empty or untrusted input before synthesis.
- Set a fixed seed for deterministic caching and tests.
- Measure on the actual target device before claiming real-time operation.

## Troubleshooting

- Missing phonemizer errors: reinstall `phonemizer` and `espeakng-loader` from `requirements.txt`.
- Unexpected pronunciation: spell out ambiguous abbreviations or add context; the release frontend is English-only.
- Flat or unstable delivery: keep `variation` near the default before changing speed.
- Long pauses: inspect source punctuation because chunk boundaries intentionally follow punctuation.
- Slow first request: model loading is measured separately from warm synthesis and should not happen per utterance.