--- title: Austrian Dialect Text-to-Speech emoji: 🗣️ colorFrom: blue colorTo: purple sdk: gradio sdk_version: 4.44.1 app_file: app.py pinned: false --- # Austrian Dialect Text-to-Speech Synthesis A neural text-to-speech (TTS) interface demonstrating FastSpeech 2 with dialect embeddings for Austrian German dialects and standard German. ## Overview This demo showcases neural speech synthesis for multiple Austrian dialects using: - **FastSpeech 2** as the TTS backbone - **IMS-Toucan** adaptation with dialect embeddings - **VoxLingua107 XLS-R** for multilingual speaker embeddings - **Standard German grapheme-to-phoneme conversion** with dialect support Supported dialects: - Standard Austrian German - Viennese (Wien) - Goisern (Upper Austrian) - Innervillgraten (Tyrolean) ## Features - **Multiple Speakers**: 12 different speaker profiles across all dialects - **Dialect Selection**: Switch between standard and dialect embeddings - **Interpolation**: Create custom dialect blends with 50/50 interpolations - **German-only Interface**: Optimized for German text input ## Usage 1. Enter German text in the text input 2. Select a speaker from the dropdown 3. Choose a dialect/standard embedding 4. Optionally select an interpolation option 5. Click "Submit" to generate audio ## Examples - "Der Nordwind und die Sonne stritten sich, wer von ihnen wohl der Stärkere wäre." (Standard) - "Grüß Gott, wie geht es dir?" (Viennese) ## Technical Details ### Model Architecture - **TTS Model**: FastSpeech 2 - **Vocoder**: Avocodo - **Phonemizer**: espeak with Austrian dialect support - **Embeddings**: XLS-R 300M wav2vec features for speaker/dialect conditioning ### Requirements - Python 3.8+ - PyTorch 1.10.2 - Gradio 3.44.3 - All dependencies listed in `requirements.txt` ### File Structure ``` . ├── app.py # Main Gradio interface ├── requirements.txt # Python dependencies ├── InferenceInterfaces/ # TTS inference implementations ├── Preprocessing/ # Text preprocessing and embeddings │ └── wav2vec_embeddings/ # Pre-computed dialect embeddings ├── Utility/ # Utility functions │ └── example_wavs/ # Speaker reference audio files ├── Models/ # Pre-trained model files └── audios/ # Output audio directory ``` ## References [1] Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, "Fastspeech 2: Fast and high-quality end-to-end text to speech," in *ICLR 2021 - 9th International Conference on Learning Representations*, 2021. [2] Florian Lux and Bilal Ghauri. IMS-Toucan: An Open-Source Framework for Training and Deploying Speech Synthesis Models. *Proc. INTERSPEECH 2023*, 2023. [3] https://huggingface.co/TalTechNLP/voxlingua107-xls-r-300m-wav2vec [4] L. Gutscher, M. Pucher, and V. García, "Neural Speech Synthesis for Austrian Dialects with Standard German Grapheme-to-Phoneme Conversion and Dialect Embeddings," in *Proc. 2nd Annual Meeting of the ELRA/ISCA SIG on Under-resourced Languages (SIGUL 2023)*, 2023. [Paper](https://www.isca-archive.org/sigul_2023/gutscher23_sigul.html) ## Authors - [Lorenz Gutscher](mailto:lorenz.gutscher@ofai.at) - OFAI Vienna - [Michael Pucher](mailto:michael.pucher@ofai.at) - OFAI Vienna ## Citation If you use this work, please cite: ```bibtex @inproceedings{gutscher2023neural, title={Neural Speech Synthesis for Austrian Dialects with Standard German Grapheme-to-Phoneme Conversion and Dialect Embeddings}, author={Gutscher, Lorenz and Pucher, Michael and García, Víctor}, booktitle={Proc. 2nd Annual Meeting of the ELRA/ISCA SIG on Under-resourced Languages (SIGUL 2023)}, year={2023} } ``` ## License Please refer to the LICENSE file for licensing information. ## Note on Data Text input and audio output are processed temporarily during use. Please see the security notice in the interface for details. --- For more information, visit the [Austrian Institute for Artificial Intelligence (OFAI)](https://www.ofai.at/)