Multimodal emotion (face + voice + Traditional Chinese text) for Reachy Mini

This repo holds three small emotion classifiers plus a late-fusion rule. They run on CPU with ONNX Runtime and need no PyTorch, and they are sized for the Raspberry Pi CM4 in a Reachy Mini Wireless. The robot app that uses them is the Space pearlyjam21/reachy_mini_multimodal_emotion.

Output labels (fusion order): neutral, happy, sad, angry, surprise, fear, disgust.

Branch File Model Input Size
Face face/fer_mobilenetv3_fp32.onnx MobileNetV3-Large, FER2013 images with FER+ soft labels 112×112 grayscale face crop (YuNet, 10% padding) replicated to RGB, ImageNet norm 17 MB
Face detector face/face_detection_yunet_2023mar.onnx YuNet (OpenCV Zoo) BGR frame 0.2 MB
Voice speech/ser_ecapa_v3.onnx compact ECAPA-TDNN student (1.9 M params), distilled 2 s, 16 kHz, 80-band log-mel, per-utterance norm 7.6 MB
Words text/text_bert_zh_int8.onnx + text/tokenizer.json bert-base-chinese fine-tune, dynamic int8 Traditional Chinese text, ≤128 tokens 98 MB
Utterance VAD asr/silero_vad.onnx Silero VAD (sherpa-onnx build) 16 kHz 0.6 MB

Speech-to-text is not stored here. The app downloads SenseVoice-Small int8 from its original repo, because it has its own license (FunASR model license).

config.json is the contract used by the code. It holds each branch's label order, file paths, preprocessing, temperatures and fusion weights.

How fusion works

  1. Each branch's logits are reordered into the fusion label order (config.json → branch_labels).
  2. Each branch is calibrated as softmax(logits / T_branch).
  3. The fused distribution is p ∝ exp(Σ w_b · log p_b), a weighted log-linear pool.
  4. Fusion runs once per spoken sentence (silero VAD start/end). The face is averaged over the frames seen while the sentence was spoken. The sum runs only over the branches that have evidence, with weights renormalised. No face, an English transcript (the text model is Chinese-only) or a sentence too short to score drops that branch. A missing branch is never counted as "neutral".
Parameter Value Where it comes from
T_face 6.46 fitted on CREMA-D validation video (YuNet crops, 5 fps)
w_face : w_speech 0.65 : 0.35 tuned on CREMA-D validation, using the v2 speech student
T_speech 1.0 not refitted for v3 (the v2 fit was 0.86)
T_text, w_text 1.0, 0.35 unvalidated defaults: no dataset aligns this text with face and voice

Evaluation

All numbers below are for single branches, or for voice + face only. No evaluation of all three combined exists yet.

Face. FER+ PublicTest (majority label; this split also selected the checkpoint, so the result is slightly optimistic): accuracy 83.6%, macro-F1 0.754. On 496 real-photo face crops with 10% padding (EmotionNet6 test, 6 classes): accuracy 59.5%, macro-F1 0.607. With 30% padding accuracy drops to 47.0%, which is why 10% is used. The int8 export of this model was rejected because its FER+ macro-F1 fell to 0.337.

Voice (v3). Unweighted average recall (UAR), one clip at a time, on held-out speakers:

ESD Mandarin (spk 0010) ESD English (0011–12) CREMA-D (9 actors) MELD test (chance 14.3%)
70.1 56.6 67.0 23.5

Words. Held-out split of Chinese-MEDD plus supplementary fear rows (n = 632, seed 42):

  • fp32: accuracy 91.9%, macro-F1 0.907.
  • int8 ONNX (shipped): accuracy 91.5%, macro-F1 0.903, agreeing with fp32 on 97.3% of rows.
  • These are human-written texts; ASR errors are not included (see reports/text_int8_eval.json).

Late fusion, voice (v2) + face, test UAR (reports/late_fusion_*):

voice only face only fused
CREMA-D (acted) 55.9 30.3 59.2
MELD (TV dialogue) 15.9 16.0 17.0

On natural conversation (MELD) the models are close to chance. Treat the output as a soft social cue, not a measurement.

Known limitations

  • Calm, quiet real voices are often classified as sad by the voice model. Neutral clips from MELD come out as happy under v3.
  • On talking-face video the face model mostly predicts neutral or happy. It was trained on posed 48×48 grayscale photos.
  • The text branch is Traditional Chinese only. Simplified output from ASR is converted with chinese-converter. The supplementary fear data is machine-generated and template-based.
  • The text branch has not been tested on real ASR transcripts. Speech-recognition errors can reverse its meaning, for example by dropping a negation.
  • The robot app's reaction thresholds were calibrated for the voice-only v2 app, not for fused outputs.
  • None of the models identifies the speaker. The robot assumes the loudest voice and the largest face belong to the same person.

Usage (PC or robot)

pip install "git+https://huggingface.co/spaces/pearlyjam21/reachy_mini_multimodal_emotion"  # private Space: run `hf auth login` first
from reachy_mini_multimodal_emotion.engine.pipeline import MultimodalPipeline

pipe = MultimodalPipeline.from_pretrained()     # downloads this repo + SenseVoice
pipe.push_frame(bgr_frame)                      # a few times per second
for sentence in pipe.push_audio(mono_16k_chunk):  # returns a sentence when the speaker pauses
    r = pipe.analyze(sentence)                  # voice + face during the sentence + words
    print(r.transcript, r.fused.top, r.fused.confidence, r.fused.used)

Provenance and licenses

These components carry different licenses. Taken together, use them for research and education only. See LICENSE.md.

Component Built from License notes
Face FER2013 images, FER+ labels (MIT), torchvision ImageNet weights (BSD-3) FER2013 images come from a Kaggle/ICML research challenge
YuNet OpenCV Zoo MIT
Voice v3 distilled from MERaLiON-SER-v1 on ESD, CREMA-D, MELD MERaLiON public licence; ESD is for non-commercial research; CREMA-D is ODbL; MELD is GPL-3.0
Words pubinfo/bert-base-chinese-cls2.0 (MIT), Chinese-MEDD (MIT), TIX007/chinese-sentiment (MIT) MIT
Silero VAD snakers4/silero-vad MIT
SenseVoice (downloaded separately) FunAudioLLM/SenseVoice FunASR model license
Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using pearlyjam21/multimodal-emotion-zh 1