--- license: mit language: - ja pipeline_tag: text-to-speech tags: - text-to-speech - speech - tts - character-image base_model: - Aratako/Irodori-TTS-500M-v2 - timm/vit_base_patch16_siglip_512.v2_webli --- # Irodori-TTS-500M-v2-Character-Voice-SigLIP [![Project Page](https://img.shields.io/badge/Project_page-orange)](https://character-voice-control.p1atdev.workers.dev/) ![arXiv](https://img.shields.io/badge/arXiv-coming_soon-b31b1b?logo=arxiv) [![GitHub](https://img.shields.io/badge/GitHub-Irodori_Character_Voice-black?logo=github)](https://github.com/p1atdev/Irodori-Character-Voice) [![Demo](https://img.shields.io/badge/Demo-HuggingFace%20Space-blue)](https://huggingface.co/spaces/p1atdev/Irodori-TTS-500M-v2-Character-Voice-Demo) ## Model Description **Irodori-TTS-500M-v2-Character-Voice-SigLIP** is a Japanese TTS model based on [Aratako/Irodori-TTS-500M-v2](https://huggingface.co/Aratako/Irodori-TTS-500M-v2). This model synthesizes speech in a specific character's voice by using the **character's image** as a condition. By using encoded features from a character image as the conditioning signal instead of reference audio or voice captions, it enables zero-shot speech synthesis with a voice that matches the character's atmosphere. This SigLIP variant uses [SigLIP-v2-B/16-512](https://huggingface.co/timm/vit_base_patch16_siglip_512.v2_webli) as the image encoder. Another model in the same Character Voice family, [Irodori-TTS-500M-v2-Character-Voice-Tagger](https://huggingface.co/p1atdev/Irodori-TTS-500M-v2-Character-Voice-Tagger), uses a wd-tagger-based image encoder instead. ## Samples > 遠くで鳴る夕暮れの鐘が、一日の終わりを告げている。家々の窓には、ぽつりぽつりと暖かな灯りがともり始めた。 | Character Image | Generated Audio | | - | - | | |