tzustu's picture
Claude Opus 5.5
0.3.6: dashboard in Traditional Chinese
8ff960b
|
Raw History Blame Contribute Delete
5.05 kB
metadata
title: Reachy Mini Multimodal Emotion
emoji: 🙂
colorFrom: indigo
colorTo: pink
sdk: static
pinned: false
short_description: Senses your emotion, moves and talks back (zh + en)
tags:
  - reachy_mini
  - reachy_mini_python_app
models:
  - pearlyjam21/multimodal-emotion-zh
  - csukuangfj/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17
  - csukuangfj/matcha-icefall-zh-en
  - csukuangfj/vits-piper-en_US-lessac-low

Reachy Mini Multimodal Emotion

Talk to Reachy Mini in Mandarin or English, one sentence at a time. It senses how you feel from your face, voice and words. It answers at once with a move, and then talks back with a short, empathetic reply.

Step What the robot does
waiting its head follows your face
listening you start speaking: the antennas perk up
thinking you pause (0.5 s): the last part of the sentence is transcribed and the sentence is scored
reacting an empathetic move starts right away. A free LLM writes a reply, which the robot's own voice speaks phrase by phrase

Everything runs on the robot except the LLM:

  • Speech-to-text: SenseVoice-Small int8 (sherpa-onnx), accurate for Mandarin and English on the robot's far-field microphone. Long sentences are transcribed in pieces while you speak (cut at short breaks of 0.2 s), so only the last piece is left when you stop and long sentences don't take much longer than short ones. MM_EMOTION_EARLY_ASR=0 waits for the whole sentence instead. A faster streaming Zipformer is available with MM_EMOTION_ASR=streaming, but it was far less accurate on the robot.
  • Voice: Matcha-TTS zh+en with a 16 kHz Vocos vocoder for Chinese replies. English replies use Piper en_US-lessac-low (16 kHz), a native US English voice, since 0.3.4: the Matcha voice is Chinese-first and gave English a Chinese accent. Both take 0.05–0.09 s of compute per second of speech on a laptop.
  • Emotion: the face, voice and Traditional Chinese text models from pearlyjam21/multimodal-emotion-zh. The text model is only used for Chinese sentences.

The LLM runs on free OpenRouter models: the fastest measured free chat models first, then the random openrouter/free route.

It always answers. If the free model hasn't started replying within the deadline (3 s by default), or there's no key, no network, or a rate limit, the robot says a short prepared line that matches your emotion.

Setup

  1. Install the app and open the dashboard at http://reachy-mini.local:8042.
  2. Paste an OpenRouter API key into Conversation. It is stored on the robot only (~/.config/reachy_mini_multimodal_emotion/openrouter.key) and never shown again. You can also set OPENROUTER_API_KEY instead.
  3. Free models allow 50 requests per day, or 1,000 per day once the account has bought $10 of credits. Each sentence you say uses one request.

On first start the robot downloads about 500 MB of models. Downloads retry automatically if the network drops.

Dashboard

  • The current step, a microphone meter, and the live transcript while you speak.
  • For each sentence: what it heard, the combined emotion, what each modality said, what the robot replied (and which model wrote it), the move it played, and the timings: emotion ready, the LLM's first word, and when the robot started speaking.
  • Controls: talking on/off, the reply deadline, moves, face following, which modalities to use, and the fusion weights.

Measured on a laptop in simulation: speech starts about 2.1–2.7 s after you stop talking, including the 0.5 s pause needed to know you've finished. On the Pi it is slower; the dashboard shows the real numbers. Since 0.3.3 an 11 s sentence is ready about 0.1 s after the pause on the laptop, instead of 0.4 s. Leaked model reasoning ("We need to respond…") is detected and never spoken. Since 0.3.5 every phrase is checked, not just the first, so a reply that slips into reasoning ("Okay, let's unpack this. The user is…") stops there. Leaked reasoning is also kept out of the conversation history, so one bad reply can't teach the model to keep doing it.

Dashboard language: Traditional Chinese (since 0.3.6). The robot's logs stay in English.

Simulation: set MM_EMOTION_WAV=/path/to/16k.wav to replay a file instead of the microphone, MM_EMOTION_LOCAL_DIR to use local emotion models, and MM_EMOTION_ASR=streaming for the faster, less accurate streaming speech-to-text.

The emotion output is a soft social cue, not a measurement. The models are near chance on natural conversation; see the model card. The Matcha voice's upstream license isn't stated in its repository. The English voice is trained on the Lessac Blizzard 2013 data; check that license before commercial use.