pearlyjam21's picture
Reachy Mini multimodal emotion app
d626f8b verified
|
Raw History Blame
1.98 kB
metadata
title: Reachy Mini Multimodal Emotion
emoji: 🙂
colorFrom: indigo
colorTo: pink
sdk: static
pinned: false
short_description: Face + voice + Chinese words emotion, empathetic reactions
tags:
  - reachy_mini
  - reachy_mini_python_app
models:
  - pearlyjam21/multimodal-emotion-zh
  - csukuangfj/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17

Reachy Mini Multimodal Emotion

This app lets Reachy Mini sense emotion from three sources and react with an empathetic move from pollen-robotics/reachy-mini-emotions-library:

  • Face: the camera runs at about 5 fps; YuNet finds the face and a FER+ MobileNetV3 classifies it.
  • Voice: the microphone feeds a 2 s window every 0.2 s to a compact ECAPA-TDNN student.
  • Words: silero VAD splits speech into utterances, SenseVoice-Small transcribes them, and a Traditional Chinese BERT classifies the text.

The three results are combined by calibrated late fusion (weighted log-linear pooling). A modality with no evidence is left out rather than counted as neutral: no face, silence, English speech (the text model is Chinese-only), or a transcript that is too old. Everything runs on the Reachy Mini Wireless CPU with ONNX Runtime, and no PyTorch is needed.

Models download from pearlyjam21/multimodal-emotion-zh on first start (about 360 MB). Its model card lists the metrics, licenses and known limitations.

Dashboard: http://reachy-mini.local:8042 shows fused and per-modality probabilities, the transcript, fusion-weight sliders, modality on/off switches and test reactions.

Simulation: set MM_EMOTION_WAV=/path/to/16k.wav to loop a file instead of the microphone. Set MM_EMOTION_LOCAL_DIR to load the models from a local folder instead of the Hub.

The output is a soft social cue, not a measurement. On natural conversation the individual models are close to chance; see the model card.