--- title: Reachy Mini Multimodal Emotion emoji: 🙂 colorFrom: indigo colorTo: pink sdk: static pinned: false short_description: Senses your emotion, moves and talks back (zh + en) tags: - reachy_mini - reachy_mini_python_app models: - pearlyjam21/multimodal-emotion-zh - csukuangfj/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17 - csukuangfj/matcha-icefall-zh-en --- # Reachy Mini Multimodal Emotion Talk to Reachy Mini in **Mandarin or English**, one sentence at a time. It senses how you feel from your **face**, **voice** and **words**. It answers at once with a **move**, and then **talks back** with a short, empathetic reply. | Step | What the robot does | |---|---| | waiting | its head follows your face | | listening | you start speaking: the antennas perk up | | thinking | you pause (0.5 s): SenseVoice transcribes the sentence and it is scored | | reacting | an empathetic move starts right away. A free LLM writes a reply, which the robot's own voice speaks phrase by phrase | **Everything runs on the robot except the LLM:** - **Speech-to-text:** SenseVoice-Small int8 (sherpa-onnx), accurate for Mandarin and English on the robot's far-field microphone. It runs after the sentence ends. A faster streaming Zipformer is available with `MM_EMOTION_ASR=streaming`, but it was far less accurate on the robot. - **Voice:** Matcha-TTS zh+en with a 16 kHz Vocos vocoder. - **Emotion:** the face, voice and Traditional Chinese text models from [`pearlyjam21/multimodal-emotion-zh`](https://huggingface.co/pearlyjam21/multimodal-emotion-zh). The text model is only used for Chinese sentences. The LLM runs on free OpenRouter models: the fastest measured free chat models first, then the random `openrouter/free` route. **It always answers.** If the free model hasn't started replying within the deadline (3 s by default), or there's no key, no network, or a rate limit, the robot says a short prepared line that matches your emotion. ## Setup 1. Install the app and open the dashboard at `http://reachy-mini.local:8042`. 2. Paste an [OpenRouter](https://openrouter.ai) API key into **Conversation**. It is stored on the robot only (`~/.config/reachy_mini_multimodal_emotion/openrouter.key`) and never shown again. You can also set `OPENROUTER_API_KEY` instead. 3. Free models allow 50 requests per day, or 1,000 per day once the account has bought $10 of credits. Each sentence you say uses one request. On first start the robot downloads about 500 MB of models. Downloads retry automatically if the network drops. ## Dashboard - The current step, a microphone meter, and the live transcript while you speak. - For each sentence: what it heard, the combined emotion, what each modality said, what the robot replied (and which model wrote it), the move it played, and the timings: emotion ready, the LLM's first word, and when the robot started speaking. - Controls: talking on/off, the reply deadline, moves, face following, which modalities to use, and the fusion weights. **Measured on a laptop in simulation:** speech starts about 2.1–2.7 s after you stop talking, including the 0.5 s pause needed to know you've finished. Expect roughly 3–5 s on the Pi, since SenseVoice needs 1.5–3 s there; the dashboard shows the real numbers. Leaked model reasoning ("We need to respond…") is detected and never spoken. **Simulation:** set `MM_EMOTION_WAV=/path/to/16k.wav` to replay a file instead of the microphone, `MM_EMOTION_LOCAL_DIR` to use local emotion models, and `MM_EMOTION_ASR=streaming` for the faster, less accurate streaming speech-to-text. The emotion output is a soft social cue, not a measurement. The models are near chance on natural conversation; see the model card. The Matcha voice's upstream license isn't stated in its repository.