clem's picture
clem HF Staff
Document transcript cleanup rollback
326692b verified
|
Raw History Blame Contribute Delete
9.29 kB
metadata
title: Miami Reachy Receptionist
emoji: 🤖
colorFrom: yellow
colorTo: pink
sdk: static
pinned: false
short_description: Reachy Mini receptionist for the Miami Hugging Face office
tags:
  - reachy_mini
  - reachy_mini_python_app

Miami Reachy Receptionist

A Reachy Mini app that acts as the receptionist for the Miami Hugging Face office.

What it does

  • Looks around while idle.
  • Detects an arriving visitor with a stronger face pipeline: OpenCV YuNet DNN face detector when available, plus Haar frontal/profile fallbacks.
  • Smoothly locks onto and follows the visitor's face with ReachyMini.look_at_image(...).
  • Keeps face lock briefly through missed frames to reduce jitter/dropout.
  • Can optionally answer spoken questions before a face is detected, but this is off by default to avoid accidental engagement.
  • When the visitor is close enough, greets them:

    Hi, welcome to the Miami Hugging Face office, nice to meet you.

  • Asks:
    1. What's your name?
    2. Who are you here to see?
    3. What's the reason for your visit?
  • Records and transcribes the visitor's answers.
  • Uses smoother flexible dialogue that can collect name / host / reason in a natural conversation.
  • Moves its head and antennas while speaking so the interaction feels more alive.
  • Shows Clem a live receptionist dashboard with transcript, visit summary, status, settings, detector status, and recent events.

Install on Reachy Mini

From the Reachy Mini dashboard, install this Space, or use the REST API:

curl -X POST http://reachy-mini.local:8000/api/apps/install \
  -H "Content-Type: application/json" \
  -d '{"url": "https://huggingface.co/spaces/clem/miami-reachy-receptionist"}'

Then start it from the dashboard. Clem can open the app UI at:

http://reachy-mini.local:8042

For Reachy Mini Lite on localhost, use:

http://localhost:8042

OpenAI API key setup from the app UI

The dashboard has a Settings section with an OpenAI API key field.

  1. Start the receptionist app from the Reachy Mini Desktop App / dashboard.
  2. Open the app UI:
    • Wireless: http://reachy-mini.local:8042
    • Lite/local: http://localhost:8042
  3. Paste the OpenAI key in OpenAI API key.
  4. Click Save settings.

The key is stored locally on that Reachy Mini / desktop app instance in a local config file, not in this public Space. Leave the key field blank when saving other settings to keep the existing saved key. Use Clear OpenAI key to remove it.

With the key set, the app enables:

  • speech-to-text transcription using whisper-1
  • spoken TTS greeting/questions using tts-1
  • optional passive Q&A using gpt-4o-mini if enabled in Settings

By default, Reachy starts the receptionist flow only after a stable, close face is detected. Passive no-face Q&A is available as a setting but is off by default to avoid accidental engagement.

Settings tips

If Reachy still struggles to detect faces:

  • Keep Face score threshold around 0.65 to avoid false positives; lower only if it misses real faces.
  • Greeting mode defaults to When face is confirmed, so Reachy starts greeting as soon as a stable face is detected.
  • Switch Greeting mode to Only when close if you want distance-gated greetings again.
  • Confirmation seconds defaults to 0.5, so it should greet quickly while still filtering one-frame false positives.
  • Keep the visitor well lit and roughly in front of the camera.
  • Make sure the dashboard shows detector yunet+haar; if it shows only haar, the app could not download/use the YuNet model and will use the fallback detector.

Useful settings:

RECEPTIONIST_OFFICE_NAME="Miami Hugging Face office"
RECEPTIONIST_HOST_NAME="Clem"
RECEPTIONIST_GREETING_MODE=face
RECEPTIONIST_NEAR_FACE_RATIO=0.10
RECEPTIONIST_FACE_SCORE_THRESHOLD=0.65
RECEPTIONIST_CONFIRMATION_SECONDS=0.5
RECEPTIONIST_MIN_TRACK_FACE_RATIO=0.012
RECEPTIONIST_PASSIVE_CONVERSATION=0
RECEPTIONIST_TALK_MOTION=1
RECEPTIONIST_FLEXIBLE_DIALOGUE=1
RECEPTIONIST_PASSIVE_LISTEN_SECONDS=4
RECEPTIONIST_GREETING_COOLDOWN_SECONDS=90
RECEPTIONIST_ANSWER_SECONDS=7

Advanced model/voice overrides can still be set in the environment if needed:

RECEPTIONIST_TTS_MODEL=tts-1
RECEPTIONIST_TTS_VOICE=alloy
RECEPTIONIST_STT_MODEL=whisper-1
RECEPTIONIST_CHAT_MODEL=gpt-4o-mini

Local development

python -m venv .venv
source .venv/bin/activate
pip install -e .
python -m miami_reachy_receptionist.main

Make sure the Reachy Mini daemon is running first.

Notes

This app follows the Reachy Mini Python app contract: it exposes a ReachyMiniApp entry point under the reachy_mini_apps group and implements run(reachy_mini, stop_event).

Conversation smoothness

The dashboard includes two controls enabled by default:

  • Move head and antennas while speaking: adds small synchronized motion during TTS.
  • Use smoother flexible receptionist dialogue: lets the visitor answer naturally instead of forcing one rigid question at a time. Reachy extracts name, who they are here to see, and reason from the conversation, then asks only for missing information.

Face lock during conversation

During a visit, the app now starts a small background face-lock loop. While OpenAI is generating speech, while Reachy is talking, and while it is listening, it keeps reading camera frames and calling look_at_image(...) so Reachy stays oriented toward the visitor instead of resetting to neutral. Speaking motion now mainly animates antennas/body yaw and preserves gaze on the face.

The flexible conversation path is driven by OpenAI chat with structured JSON output. Reachy uses the model to decide the next natural receptionist reply and to extract visitor name, who they are here to see, and visit reason.

Low-latency conversation

The app no longer waits the full listen window before responding. With Reply as soon as the visitor stops speaking enabled, Reachy records until it detects speech has started and then about 0.8s of silence. The Max answer listen seconds setting is only a timeout. For snappier responses, set:

Reply as soon as the visitor stops speaking: ON
Max answer listen seconds: 4
Silence before reply seconds: 0.6-0.8
Minimum utterance seconds: 0.5-0.7

Faster-than-before responses

The app now uses two latency reductions:

  • Faster OpenAI transcription model: defaults STT to gpt-4o-mini-transcribe instead of whisper-1.
  • Streaming TTS: starts pushing PCM audio chunks to Reachy as soon as OpenAI starts returning speech, instead of waiting for the whole WAV to be generated first.

This still is not as low-latency as the official Reachy Mini Conversation app's full OpenAI Realtime WebSocket pipeline, but it should feel much snappier than request/response WAV mode. For the lowest latency in this app, use:

Reply as soon as the visitor stops speaking: ON
Start speaking as soon as audio streams in: ON
Use faster OpenAI transcription model: ON
Silence before reply seconds: 0.3-0.5
Max answer listen seconds: 3-4

OpenAI Realtime mode

The receptionist app now defaults to OpenAI Realtime for the live visitor conversation, matching the low-latency architecture used by the official Reachy Mini Conversation app:

  • microphone audio is streamed continuously to OpenAI Realtime,
  • OpenAI server VAD detects turns,
  • response audio streams back chunk-by-chunk,
  • Reachy keeps the background face-lock loop running while listening and speaking.

In Settings, keep:

Conversation backend: OpenAI Realtime (lowest latency)
Use smoother flexible receptionist dialogue: ON
Move head and antennas while speaking: ON

If Realtime fails for any reason, the app falls back to the turn-based pipeline.

Editing the conversational system prompt

The Settings panel exposes Conversational agent system prompt. This prompt is used for the OpenAI Realtime receptionist session and the turn-based fallback. Edit it to change Reachy's personality, office policies, check-in questions, or escalation behavior.

Important: keep an instruction that tells Reachy to end completed check-ins with the exact phrase check-in complete; the app uses that phrase to know when the Realtime visit is done.

Improving visitor speech understanding

The Settings panel exposes speech recognition controls:

  • Speech transcription hints: add names and office-specific words, e.g. Clem, Hugging Face, Miami office, demo, interview, meeting.
  • Transcription language: set to en for English, or another ISO code if needed.
  • Mic gain: boosts quiet visitor speech before transcription. Start around 1.8; reduce if clipping/noisy.
  • Noise gate: removes very low background noise. Start around 0.003; increase slightly in noisy rooms.

The app now also chooses the louder microphone channel instead of averaging the mic array, which avoids phase cancellation and usually improves voice clarity.

Transcript cleanup rollback

The extra second OpenAI pass that tried to clean noisy transcripts has been removed because it could interfere with the live conversation. Speech understanding now uses the Realtime API transcription/hints directly, without a second chat-completion cleanup step.