You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

PersonaPlex 7B v1 β€” 4-bit NF4 Quantized (bitsandbytes)

This is a 4-bit NF4 quantized version of nvidia/personaplex-7b-v1 using bitsandbytes.

PersonaPlex is a real-time, full-duplex speech-to-speech conversational model with persona control through text-based role prompts and audio-based voice conditioning.

Why Quantize?

The original model requires ~14GB VRAM (bf16), which exceeds consumer GPUs like the RTX 4070 (12GB). This 4-bit quantized version:

Original (bf16) Quantized (NF4)
VRAM ~14 GiB ~9.6 GiB
GPU A100 / H100 RTX 4070+ (12GB)
torch.compile Yes Yes
CUDA graphs Yes Yes

What's Quantized?

Only the main transformer's linear layers (attention projections + gating FFN) are quantized to 4-bit NF4. The following are kept in bf16 for quality:

  • Mimi audio encoder/decoder
  • Depformer (depth transformer)
  • Embedding layers
  • Output heads

Quick Start

Prerequisites

  1. Accept the PersonaPlex license (required for the base model assets)
  2. Set your HuggingFace token:
export HF_TOKEN=<YOUR_TOKEN>

Installation

git clone https://huggingface.co/brianmatzelle/personaplex-7b-v1-bnb-4bit
cd personaplex-7b-v1-bnb-4bit
pip install moshi/.
pip install bitsandbytes

Run (Live Server)

SSL_DIR=$(mktemp -d)
python -m moshi.server --ssl "$SSL_DIR" --quantize-4bit

Then open https://localhost:8998 in your browser.

Only one browser/client WebSocket session is accepted at a time; extra connections get close code 1013 / SESSION_BUSY.

For 8-bit inference (higher VRAM than 4-bit, typically better quality retention):

SSL_DIR=$(mktemp -d)
python -m moshi.server --ssl "$SSL_DIR" --quantize-8bit

Run (FastAPI Production Server β€” API-only by default)

PersonaPlex is delivered as an API-first product. The supported public integration surface is the documented WebSocket + REST API; the bundled browser demo UI is disabled by default and must be opted into explicitly for internal demos.

python -m moshi.fastapi_server --quantize-4bit --host 0.0.0.0 --port 8998

Public surface (only paths an integrator should depend on):

  • WS /api/chat β€” binary voice/text streaming protocol (canonical buyer path)
  • GET / β€” JSON discovery document describing this API surface
  • GET /health β€” basic liveness
  • GET /ready β€” runtime + GPU readiness
  • GET /config β€” effective runtime config (includes webui_enabled, public_surface)
  • GET /metrics β€” Prometheus metrics

Notes:

  • Hitting /api/chat in a browser tab is expected to fail; it is a WebSocket endpoint.
  • GET /config reports webui_enabled: false and public_surface: "websocket+rest" in the default API-only mode.
  • The service runs one inference session at a time (exclusive slot). Each WebSocket must include voice_prompt (when the server is configured with a voice library) and may include text_prompt. Optional per-session overrides (query string): context (attention span, clamped to the server's --context cap), audio_temperature (alias audio_temp), audio_top_k (alias audio_topk), text_temperature (alias text_temp), text_top_k (alias text_topk). Values are restored after the session ends so the next client gets a clean baseline.
  • Latency tuning: server logs a session telemetry line every ~5 s with avg_step_ms, max_step_ms, rtf (real-time factor; <1.0 keeps up with audio), pcm_backlog_ms, and the LM lm_offset/KV-cache fill. If rtf rises above 1.0 you'll also see session falling behind real-time warnings and clients receive a BEHIND_REALTIME JSON diagnostic. Per-step attention cost grows linearly with the KV-cache fill until lm_offset reaches the server's context (then stabilizes), so for latency-sensitive deployments lower the per-session ?context= (or the server-wide --context).

Useful env overrides:

  • PERSONAPLEX_PORT
  • PERSONAPLEX_HOST
  • PERSONAPLEX_SESSION_IDLE_TIMEOUT_SECONDS
  • PERSONAPLEX_SESSION_ACQUIRE_WAIT_SECONDS (default 60) β€” how long a second WebSocket may wait for the exclusive session slot. After this wait, close code 1013 / SESSION_BUSY.
  • PERSONAPLEX_ENABLE_WEBUI β€” set to 1 to mount the bundled browser demo UI at / (off by default).
  • PERSONAPLEX_WEBUI_DIST β€” optional path to the static UI folder used when the demo UI is enabled (defaults to ./webui/dist).
  • PERSONAPLEX_UVICORN_FULL_WS_URL β€” set to 1 to restore full WebSocket URLs in uvicorn access/error logs (default is shortened; JSON logs still show applied sampling).
  • PERSONAPLEX_LOG_LEVEL

Test as a client (no UI required)

Buyers integrate against the WebSocket protocol, not against any HTML/JS the server might ship. The repo includes a thin reference client that proves a deployment is healthy end-to-end without ever opening a browser.

  1. Quick REST sanity:
curl http://<host>:<port>/             # discovery JSON (API-only mode)
curl http://<host>:<port>/health
curl http://<host>:<port>/ready
curl http://<host>:<port>/config
  1. Stateless pre-flight check (no session slot consumed):
curl 'http://<host>:<port>/api/probe?voice_prompt=NATF2.pt' \
     -H 'x-internal-auth: <token-if-configured>'
# Returns 200 + {ready: true, ...} when auth, voice_prompt, and the
# inference slot are all good; otherwise 409 with which field failed.
  1. WebSocket handshake smoke test (no audio):
# Either of these confirms the WS upgrade succeeds and (when configured)
# auth is correct. They will exit because they don't speak the binary
# protocol -- that is expected.
websocat -k 'ws://<host>:<port>/api/chat?voice_prompt=NATF2.pt'
# or
wscat -c 'ws://<host>:<port>/api/chat?voice_prompt=NATF2.pt'

Postman/wscat note: the WebSocket is a binary audio protocol. Any bad frame content -- text frame, empty binary frame, or unknown binary kind byte -- is reported as a non-fatal JSON warning text frame back from the server, and the session stays open so you can correct the next frame and continue. Example after sending the text Hello:

{"warning": "TEXT_FRAMES_NOT_SUPPORTED",
 "fatal": false,
 "message": "WS /api/chat is a binary audio protocol. Send 0x01 + Opus(24 kHz mono) frames. This text frame was ignored; the session is still open.",
 "hint": "See GET / for the frame protocol; use moshi.api_client as a reference."}

The session is only closed for real problems: auth failure (1008), capacity/setup (1013), session timeout or internal error (1011). Postman is therefore useful for connectivity, auth, and protocol-validation experiments; it cannot drive a real voice session because it cannot generate Opus audio.

  1. Postman-friendly audio in / WAV out (POST /api/voice/once β€” no Opus coding on your side):

Formats

  • WAV / FLAC / AIFF (anything libsndfile can read): send as raw body; Content-Type can be audio/wav or even application/octet-stream.
  • WebM / MP4 (the options Postman shows, e.g. audio/webm, video/webm, audio/mp4, video/mp4): set Headers β†’ Content-Type to match the file you attach. The server decodes these with ffmpeg (the ffmpeg binary must be on the server PATH, e.g. apt install ffmpeg on Ubuntu).
# curl: send a WAV, get a WAV back
curl -X POST 'http://<host>:<port>/api/voice/once?voice_prompt=NATF2.pt' \
     -H 'Content-Type: audio/wav' \
     -H 'x-internal-auth: <token-if-configured>' \
     --data-binary @assets/test/input_assistant.wav \
     -D /tmp/headers.txt \
     -o agent_reply.wav

# curl: send a WebM recording (after installing ffmpeg on the server)
curl -X POST 'http://<host>:<port>/api/voice/once?voice_prompt=NATF2.pt' \
     -H 'Content-Type: audio/webm' \
     --data-binary @recording.webm \
     -D /tmp/headers.txt \
     -o agent_reply.wav
# Transcript JSON is URL-encoded in the x-transcript response header
python -c "from urllib.parse import unquote; import re; print(unquote([l for l in open('/tmp/headers.txt') if l.lower().startswith('x-transcript:')][0].split(': ', 1)[1].strip()))"

In Postman: POST to the same URL, Body β†’ binary, pick your file. Set Headers β†’ Content-Type to audio/webm, video/webm, audio/mp4, or video/mp4 when you are not sending WAV. The response body is always a WAV (use Save Response β†’ Save to file and play locally); x-transcript holds the URL-encoded JSON transcript.

If WebM/MP4 fails with 503 / FFMPEG_NOT_AVAILABLE, install ffmpeg on the inference host. If soundfile cannot open the file and ffmpeg is missing, you get 400 / INVALID_AUDIO_BODY.

This endpoint is synchronous and shares the single-session lock with /api/chat, so it is meant for testing and batch evaluation, not for real-time concurrent traffic.

  1. Full voice session via the bundled reference client (moshi.api_client):
pip install moshi/.   # picks up websockets + soundfile

# Headless: drive a WAV file in, capture agent audio + transcript out
python -m moshi.api_client \
  --url ws://<host>:<port>/api/chat \
  --voice-prompt NATF2.pt \
  --text-prompt "You are a friendly assistant." \
  --input-wav assets/test/input_assistant.wav \
  --output-wav agent_reply.wav

# Or live mic + speaker (requires a working audio device)
python -m moshi.api_client \
  --url ws://<host>:<port>/api/chat \
  --voice-prompt NATF2.pt \
  --mic --speaker --duration 30

Streaming text tokens print to stdout as the agent talks; received audio goes to --output-wav and/or the speaker. Pass --internal-auth-token … when the deployment is configured with a shared secret, and --tenant-id … to populate the x-tenant-id header. After install, the same client is also available as personaplex-client ….

The WebSocket binary protocol is summarized in moshi/moshi/api_client.py's module docstring; that file is the single source of truth for what integrators have to implement.

  1. Standalone desktop client folder (clients/womens-helpline-desktop/) β€” a self-contained Women's Healthcare Helpline demo app you can copy to any laptop without cloning this repo. The inference stack stays remote on the server; only this folder needs to live on the tester's machine.

    What's in the folder:

    • app.py β€” single-file Tkinter desktop client (Start Call / End Call button, status pill, mm:ss timer, live transcript).
    • requirements.txt β€” 4 runtime deps (numpy, sphn, websockets, sounddevice). No torch, fastapi, or moshi.
    • README.md β€” 1-page setup + run instructions for the tester.

    On the tester's laptop:

    # copy or download just clients/womens-helpline-desktop/
    cd womens-helpline-desktop
    python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
    pip install -r requirements.txt
    # On Linux only: sudo apt install python3-tk
    
    python app.py \
      --url ws://<host>:<port>/api/chat \
      --voice-prompt NATF2.pt
      # optional: --internal-auth-token <token>
    

    The persona prompt defaults to a women's-healthcare-helpline tone (warm, non-judgmental, never diagnoses, defers to qualified doctors). The tester can override it directly in the UI before clicking Start Call, or via --text-prompt.

Productization: how to ship this to your end customers

Short answer: don't ship the raw WebSocket URL to end customers as the primary integration surface. It works for back-end-to-back-end usage and for advanced integrators, but a "click a button to talk to the agent" experience needs a thinner abstraction in front of it. The intended architecture (mirroring what Vapi, Retell, ElevenLabs, Daily, Twilio do) has four layers, in order of who owns each one:

Layer Who owns it What it does
Inference + WS API You (this repo) POST /api/voice/once, WS /api/chat, /api/probe -- the binary contract integrators can program against.
Control plane (token mint) You An auth/billing service buyers' backends call (POST /v1/sessions) to mint short-lived per-session tokens.
Client SDK (npm / pip / Swift) You Wraps WebAudio + Opus encode/decode + WS; exposes a connect(token) API with start()/stop() events.
Embeddable widget You (or buyer) One-line <script>/<iframe> drop-in for non-engineering buyers.
Customer app Customer Calls SDK or drops in widget. Never sees the raw WS contract.

For the AWS Marketplace launch in the project plan, the practical sequencing is:

  1. Now (done in this repo) β€” harden the API surface (WS + REST + probe + diagnostics).
  2. Next (Phase 2/3 of the plan) β€” build the control plane: per-tenant subscription store, ResolveCustomer / GetEntitlements, short-lived session token issuance, MeterUsage for billing. The buyer's backend exchanges their AWS Marketplace customer identifier for a session token, then hands that token to the buyer's frontend. The frontend uses the token in x-internal-auth (or a successor header you choose for the public API) when opening /api/chat.
  3. Then β€” publish a JS SDK (@personaplex/web): await PersonaPlex.connect({ token }) handles mic capture, Opus encoding, WS lifecycle, audio playback, and exposes events (onTranscript, onAudio, onConnected, onDisconnected).
  4. Optional β€” a one-line embed (<script src="https://cdn.personaplex.ai/embed.js" data-token-endpoint="...">) that drops a "Talk to agent" button into any page.

Why not skip straight to "give the customer the WebSocket URL":

  • Browsers can't easily produce Opus-encoded audio; an SDK is required.
  • You can't gate per-tenant usage, rotate tokens, or rate-limit at the WS level cleanly -- that belongs in a session-token mint endpoint.
  • Marketplace billing requires MeterUsage calls that depend on your control plane knowing when a session started/ended -- the SDK is the natural place to emit those signals.

The current repo is the right foundation. The pieces above are explicit follow-on work and live in Phases 2-4 of the project plan.

Production tuning for real-time voice latency

Real-time voice has a tight latency budget (~150–300 ms is the threshold for "feels live"). Server inference time is only one component of that budget; the rest is dominated by network RTT and audio buffering on the client. Here is the playbook we use, in order of impact.

  1. Co-locate the server with your users. Cross-continent RTT alone can be 200–400 ms. Run the inference deployment in the same AWS region as the buyer's customers. For a multi-region rollout, front the API with GeoDNS or Route 53 latency-based routing.
  2. Tune --context to the call length you actually serve. Per-step attention cost grows linearly with min(state.offset, context) until the KV cache fills. With --context 4000 at 12.5 fps the cache fills in ~320 s and worst-case attention runs over 4000 tokens. For a typical 30–60 s helpline call, set the per-session ?context= to 500–1500 β€” the cache fills in 40–120 s and the steady-state per-step GPU time is much lower. Server logs show avg_step_ms and rtf per session so you can pick the right value empirically.
  3. Watch the per-session telemetry. Every ~5 s the server emits a session telemetry line with avg_step_ms, max_step_ms, rtf (real-time factor; <1 is healthy), pcm_backlog_ms, and lm_offset. When rtf is consistently <0.5 the server is the easy part of the pipeline. When it climbs above ~0.85 you have ~15% headroom left and should lower context. The metrics are also exposed at /metrics as personaplex_lm_step_seconds and personaplex_pcm_backlog_ms.
  4. TCP_NODELAY is on by default for /api/chat connections (the server sets it after accept()). Don't put a Layer-7 proxy in front of the WebSocket that re-buffers small frames. nginx is fine if you set proxy_buffering off and proxy_request_buffering off on the WebSocket location (the bundled deploy/nginx/personaplex.conf already does).
  5. Fix client-side audio buffering. The largest single win is forcing the client's playback library to use a small buffer:
    • sounddevice.OutputStream(latency='low', blocksize=480) (Python; the bundled desktop demo already uses this)
    • WebAudio: build the SDK on top of AudioWorklet with a 128- or 256-sample worklet processor, not a <audio> tag (which buffers hundreds of ms).
    • Native iOS/Android: use the low-latency audio APIs (AVAudioEngine, AAudio) with the smallest supported buffer size; do not use AVAudioPlayer or MediaPlayer.
  6. Cap concurrency at the model. This server runs one session at a time per process (the exclusive session gate). For more concurrent callers, run multiple processes (one per GPU) behind a router that round-robins by tenant_id or by pure least-loaded; do not stack sessions on a single GPU process β€” KV-cache memory and step latency both scale with concurrent sessions and you'll lose real-time guarantees.
  7. Add a health-aware load balancer. GET /ready returns 503 until warmup completes; point your LB's health check at it so a freshly booted instance doesn't take traffic before the GPU is hot.
  8. Pin your CUDA + PyTorch versions and use CUDA graphs (already enabled via CUDAGraphed in LMGen). Driver/torch upgrades can shift per-step latency by 10–30 ms; treat them as deployment events with pre/post-deployment latency benchmarks.
  9. Observability you should add for production.
    • Structured logs (this repo emits JSON already) shipped to your log pipeline. Always include request_id, session_id, tenant_id.
    • /metrics scraped by Prometheus + Grafana with alerts on personaplex_pcm_backlog_ms p99 > 500 ms, personaplex_lm_step_seconds p99 > 70 ms, and personaplex_active_sessions stuck at 1 for >5 min (deadlocked session).
    • Per-call recording of the BEHIND_REALTIME diagnostic count.
  10. At the contract layer, version the protocol. Today the WebSocket speaks one binary protocol; the moment a buyer integrates against /api/chat, you cannot change frame shapes without a coordinated cutover. Add a ?protocol_version= query parameter and a JSON diagnostic on connect that includes the negotiated version, so future breaking changes can be additive.

If you ever see rtf < 0.5 in telemetry but the user still reports "agent feels slow", the bottleneck is not the server β€” it's RTT, the client's playback buffer, or end-pointing (the model's decision about when the user has finished speaking). Diagnose in that order.

Optional: enable the bundled demo UI (internal use only)

The repo ships webui/dist (PersonaPlex browser UI). It is not the buyer-facing integration path β€” buyers integrate against the documented WebSocket protocol. To run the demo locally for evaluation:

PERSONAPLEX_ENABLE_WEBUI=1 \
python -m moshi.fastapi_server --quantize-4bit --host 0.0.0.0 --port 8998
# or equivalently:
python -m moshi.fastapi_server --quantize-4bit --enable-webui

When enabled, the server serves webui/dist at /. If the bundle is missing locally, it is downloaded from Hugging Face (dist.tgz). For production we recommend hosting any demo UI on a separate hostname and never mixing it with the paid API surface.

Deploy Behind Nginx (TLS + WebSocket)

For public usage, terminate TLS at Nginx and proxy to the FastAPI process.

  1. Start the inference server on loopback:
python -m moshi.fastapi_server \
  --quantize-4bit \
  --host 127.0.0.1 \
  --port 8004 \
  --moshi-weight /home/user/personaplex-7b-v1-bnb-4bit/model_bnb_4bit_finetuned.pt \
  --internal-auth-token "change-me"
  1. Install Nginx config:
sudo cp deploy/nginx/personaplex.conf /etc/nginx/sites-available/personaplex
sudo ln -sf /etc/nginx/sites-available/personaplex /etc/nginx/sites-enabled/personaplex
sudo nginx -t
sudo systemctl reload nginx
  1. Edit deploy/nginx/personaplex.conf values before reload:
  • replace api.example.com with your real domain
  • replace cert paths with your real certificate files
  • ensure upstream personaplex_fastapi points at your FastAPI listen address (default 127.0.0.1:8004)
  1. Customer-facing URLs (API-only, no UI):
  • Streaming API (WebSocket): wss://api.example.com/api/chat?text_prompt=...&voice_prompt=...
  • API discovery: https://api.example.com/
  • REST health/readiness/config: https://api.example.com/health, /ready, /config
  • Prometheus metrics (operator-only): https://api.example.com/metrics

The shipped deploy/nginx/personaplex.conf proxies only the API paths and returns 404 for everything else. The bundled browser demo UI is intentionally not exposed via this config; if you want to run a demo, host it on a separate hostname/port and start the backend with --enable-webui (see above).

Notes:

  • GET /api/chat in a browser tab is expected to fail; it is a WebSocket endpoint.
  • If using an IP instead of a domain, browser TLS and audio APIs can behave inconsistently. Use a domain + valid cert for production.

Run (Offline Evaluation)

python -m moshi.offline \
  --voice-prompt "NATF2.pt" \
  --input-wav "assets/test/input_assistant.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json" \
  --quantize-4bit

8-bit offline variant:

python -m moshi.offline \
  --voice-prompt "NATF2.pt" \
  --input-wav "assets/test/input_assistant.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json" \
  --quantize-8bit

Using Pre-Quantized Weights

This repo includes pre-quantized weights (model_bnb_4bit.pt) so you don't need the full 16.7GB download. To use them, pass --moshi-weight model_bnb_4bit.pt along with --quantize-4bit. The loader auto-detects the pre-quantized format and skips re-quantization.

Changes from Base Model

This repo includes a modified moshi/ package with:

  • --quantize-4bit flag for on-the-fly 4-bit NF4 quantization via bitsandbytes
  • --quantize-8bit flag for on-the-fly 8-bit int8 quantization via bitsandbytes
  • Pre-quantized checkpoint loading (auto-detected, no re-quantization needed)
  • --cpu-offload fixes for consumer GPU compatibility
  • Attention in_proj refactored as a proper nn.Module for quantization support
  • Gating forward path updated to route through quantized modules

Citation

@misc{roy2026personaplexvoicerolecontrol,
      title={PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models},
      author={Rajarshi Roy and Jonathan Raiman and Sang-gil Lee and Teodor-Dumitru Ene and Robert Kirby and Sungwon Kim and Jaehyeon Kim and Bryan Catanzaro},
      year={2026},
      eprint={2602.06053},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2602.06053},
}

License

Code is MIT licensed. Model weights are under the NVIDIA Open Model License.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for alimuhammad9/personaplex-7b-v1-BNB-4bit

Quantized
(10)
this model

Paper for alimuhammad9/personaplex-7b-v1-BNB-4bit