--- title: Miami Reachy Receptionist emoji: 🤖 colorFrom: yellow colorTo: pink sdk: static pinned: false short_description: Reachy Mini receptionist for the Miami Hugging Face office tags: - reachy_mini - reachy_mini_python_app --- # Miami Reachy Receptionist A Reachy Mini app that acts as the receptionist for the Miami Hugging Face office. ## What it does - Looks around while idle. - Detects an arriving visitor with a stronger face pipeline: OpenCV YuNet DNN face detector when available, plus Haar frontal/profile fallbacks. - Smoothly locks onto and follows the visitor's face with `ReachyMini.look_at_image(...)`. - Keeps face lock briefly through missed frames to reduce jitter/dropout. - Can optionally answer spoken questions before a face is detected, but this is off by default to avoid accidental engagement. - When the visitor is close enough, greets them: > Hi, welcome to the Miami Hugging Face office, nice to meet you. - Asks: 1. What's your name? 2. Who are you here to see? 3. What's the reason for your visit? - Records and transcribes the visitor's answers. - Uses smoother flexible dialogue that can collect name / host / reason in a natural conversation. - Moves its head and antennas while speaking so the interaction feels more alive. - Shows Clem a live receptionist dashboard with transcript, visit summary, status, settings, detector status, and recent events. ## Install on Reachy Mini From the Reachy Mini dashboard, install this Space, or use the REST API: ```bash curl -X POST http://reachy-mini.local:8000/api/apps/install \ -H "Content-Type: application/json" \ -d '{"url": "https://huggingface.co/spaces/clem/miami-reachy-receptionist"}' ``` Then start it from the dashboard. Clem can open the app UI at: ```text http://reachy-mini.local:8042 ``` For Reachy Mini Lite on localhost, use: ```text http://localhost:8042 ``` ## OpenAI API key setup from the app UI The dashboard has a **Settings** section with an **OpenAI API key** field. 1. Start the receptionist app from the Reachy Mini Desktop App / dashboard. 2. Open the app UI: - Wireless: `http://reachy-mini.local:8042` - Lite/local: `http://localhost:8042` 3. Paste the OpenAI key in **OpenAI API key**. 4. Click **Save settings**. The key is stored locally on that Reachy Mini / desktop app instance in a local config file, not in this public Space. Leave the key field blank when saving other settings to keep the existing saved key. Use **Clear OpenAI key** to remove it. With the key set, the app enables: - speech-to-text transcription using `whisper-1` - spoken TTS greeting/questions using `tts-1` - optional passive Q&A using `gpt-4o-mini` if enabled in Settings By default, Reachy starts the receptionist flow only after a stable, close face is detected. Passive no-face Q&A is available as a setting but is off by default to avoid accidental engagement. ## Settings tips If Reachy still struggles to detect faces: - Keep **Face score threshold** around `0.65` to avoid false positives; lower only if it misses real faces. - **Greeting mode** defaults to `When face is confirmed`, so Reachy starts greeting as soon as a stable face is detected. - Switch Greeting mode to `Only when close` if you want distance-gated greetings again. - **Confirmation seconds** defaults to `0.5`, so it should greet quickly while still filtering one-frame false positives. - Keep the visitor well lit and roughly in front of the camera. - Make sure the dashboard shows detector `yunet+haar`; if it shows only `haar`, the app could not download/use the YuNet model and will use the fallback detector. Useful settings: ```env RECEPTIONIST_OFFICE_NAME="Miami Hugging Face office" RECEPTIONIST_HOST_NAME="Clem" RECEPTIONIST_GREETING_MODE=face RECEPTIONIST_NEAR_FACE_RATIO=0.10 RECEPTIONIST_FACE_SCORE_THRESHOLD=0.65 RECEPTIONIST_CONFIRMATION_SECONDS=0.5 RECEPTIONIST_MIN_TRACK_FACE_RATIO=0.012 RECEPTIONIST_PASSIVE_CONVERSATION=0 RECEPTIONIST_TALK_MOTION=1 RECEPTIONIST_FLEXIBLE_DIALOGUE=1 RECEPTIONIST_PASSIVE_LISTEN_SECONDS=4 RECEPTIONIST_GREETING_COOLDOWN_SECONDS=90 RECEPTIONIST_ANSWER_SECONDS=7 ``` Advanced model/voice overrides can still be set in the environment if needed: ```env RECEPTIONIST_TTS_MODEL=tts-1 RECEPTIONIST_TTS_VOICE=alloy RECEPTIONIST_STT_MODEL=whisper-1 RECEPTIONIST_CHAT_MODEL=gpt-4o-mini ``` ## Local development ```bash python -m venv .venv source .venv/bin/activate pip install -e . python -m miami_reachy_receptionist.main ``` Make sure the Reachy Mini daemon is running first. ## Notes This app follows the Reachy Mini Python app contract: it exposes a `ReachyMiniApp` entry point under the `reachy_mini_apps` group and implements `run(reachy_mini, stop_event)`. ## Conversation smoothness The dashboard includes two controls enabled by default: - **Move head and antennas while speaking**: adds small synchronized motion during TTS. - **Use smoother flexible receptionist dialogue**: lets the visitor answer naturally instead of forcing one rigid question at a time. Reachy extracts name, who they are here to see, and reason from the conversation, then asks only for missing information. ## Face lock during conversation During a visit, the app now starts a small background face-lock loop. While OpenAI is generating speech, while Reachy is talking, and while it is listening, it keeps reading camera frames and calling `look_at_image(...)` so Reachy stays oriented toward the visitor instead of resetting to neutral. Speaking motion now mainly animates antennas/body yaw and preserves gaze on the face. The flexible conversation path is driven by OpenAI chat with structured JSON output. Reachy uses the model to decide the next natural receptionist reply and to extract visitor name, who they are here to see, and visit reason. ## Low-latency conversation The app no longer waits the full listen window before responding. With **Reply as soon as the visitor stops speaking** enabled, Reachy records until it detects speech has started and then about `0.8s` of silence. The `Max answer listen seconds` setting is only a timeout. For snappier responses, set: ```text Reply as soon as the visitor stops speaking: ON Max answer listen seconds: 4 Silence before reply seconds: 0.6-0.8 Minimum utterance seconds: 0.5-0.7 ``` ## Faster-than-before responses The app now uses two latency reductions: - **Faster OpenAI transcription model**: defaults STT to `gpt-4o-mini-transcribe` instead of `whisper-1`. - **Streaming TTS**: starts pushing PCM audio chunks to Reachy as soon as OpenAI starts returning speech, instead of waiting for the whole WAV to be generated first. This still is not as low-latency as the official Reachy Mini Conversation app's full OpenAI Realtime WebSocket pipeline, but it should feel much snappier than request/response WAV mode. For the lowest latency in this app, use: ```text Reply as soon as the visitor stops speaking: ON Start speaking as soon as audio streams in: ON Use faster OpenAI transcription model: ON Silence before reply seconds: 0.3-0.5 Max answer listen seconds: 3-4 ``` ## OpenAI Realtime mode The receptionist app now defaults to **OpenAI Realtime** for the live visitor conversation, matching the low-latency architecture used by the official Reachy Mini Conversation app: - microphone audio is streamed continuously to OpenAI Realtime, - OpenAI server VAD detects turns, - response audio streams back chunk-by-chunk, - Reachy keeps the background face-lock loop running while listening and speaking. In Settings, keep: ```text Conversation backend: OpenAI Realtime (lowest latency) Use smoother flexible receptionist dialogue: ON Move head and antennas while speaking: ON ``` If Realtime fails for any reason, the app falls back to the turn-based pipeline. ## Editing the conversational system prompt The Settings panel exposes **Conversational agent system prompt**. This prompt is used for the OpenAI Realtime receptionist session and the turn-based fallback. Edit it to change Reachy's personality, office policies, check-in questions, or escalation behavior. Important: keep an instruction that tells Reachy to end completed check-ins with the exact phrase `check-in complete`; the app uses that phrase to know when the Realtime visit is done. ## Improving visitor speech understanding The Settings panel exposes speech recognition controls: - **Speech transcription hints**: add names and office-specific words, e.g. `Clem, Hugging Face, Miami office, demo, interview, meeting`. - **Transcription language**: set to `en` for English, or another ISO code if needed. - **Mic gain**: boosts quiet visitor speech before transcription. Start around `1.8`; reduce if clipping/noisy. - **Noise gate**: removes very low background noise. Start around `0.003`; increase slightly in noisy rooms. The app now also chooses the louder microphone channel instead of averaging the mic array, which avoids phase cancellation and usually improves voice clarity. ## Transcript cleanup rollback The extra second OpenAI pass that tried to clean noisy transcripts has been removed because it could interfere with the live conversation. Speech understanding now uses the Realtime API transcription/hints directly, without a second chat-completion cleanup step.