File size: 1,982 Bytes
d626f8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
---

title: Reachy Mini Multimodal Emotion
emoji: 🙂
colorFrom: indigo
colorTo: pink
sdk: static
pinned: false
short_description: Face + voice + Chinese words emotion, empathetic reactions
tags:
 - reachy_mini
 - reachy_mini_python_app
models:
 - pearlyjam21/multimodal-emotion-zh
 - csukuangfj/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17
---


# Reachy Mini Multimodal Emotion

This app lets Reachy Mini sense emotion from three sources and react with an empathetic move from
`pollen-robotics/reachy-mini-emotions-library`:

- **Face**: the camera runs at about 5 fps; YuNet finds the face and a FER+ MobileNetV3 classifies it.
- **Voice**: the microphone feeds a 2 s window every 0.2 s to a compact ECAPA-TDNN student.
- **Words**: silero VAD splits speech into utterances, SenseVoice-Small transcribes them, and a Traditional
  Chinese BERT classifies the text.

The three results are combined by calibrated late fusion (weighted log-linear pooling). A modality with no
evidence is left out rather than counted as neutral: no face, silence, English speech (the text model is
Chinese-only), or a transcript that is too old. Everything runs on the Reachy Mini Wireless CPU with ONNX
Runtime, and no PyTorch is needed.

Models download from [`pearlyjam21/multimodal-emotion-zh`](https://huggingface.co/pearlyjam21/multimodal-emotion-zh)
on first start (about 360 MB). Its model card lists the metrics, licenses and known limitations.

**Dashboard:** `http://reachy-mini.local:8042` shows fused and per-modality probabilities, the transcript,
fusion-weight sliders, modality on/off switches and test reactions.

**Simulation:** set `MM_EMOTION_WAV=/path/to/16k.wav` to loop a file instead of the microphone. Set
`MM_EMOTION_LOCAL_DIR` to load the models from a local folder instead of the Hub.

The output is a soft social cue, not a measurement. On natural conversation the individual models are close
to chance; see the model card.