Instructions to use alimuhammad9/personaplex-7b-v1-BNB-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Moshi
How to use alimuhammad9/personaplex-7b-v1-BNB-4bit with Moshi:
# pip install moshi # Run the interactive web server python -m moshi.server --hf-repo "alimuhammad9/personaplex-7b-v1-BNB-4bit" # Then open https://localhost:8998 in your browser
# pip install moshi import torch from moshi.models import loaders # Load checkpoint info from HuggingFace checkpoint = loaders.CheckpointInfo.from_hf_repo("alimuhammad9/personaplex-7b-v1-BNB-4bit") # Load the Mimi audio codec mimi = checkpoint.get_mimi(device="cuda") mimi.set_num_codebooks(8) # Encode audio (24kHz, mono) wav = torch.randn(1, 1, 24000 * 10) # [batch, channels, samples] with torch.no_grad(): codes = mimi.encode(wav.cuda()) decoded = mimi.decode(codes) - Notebooks
- Google Colab
- Kaggle
PersonaPlex 7B v1 β 4-bit NF4 Quantized (bitsandbytes)
This is a 4-bit NF4 quantized version of nvidia/personaplex-7b-v1 using bitsandbytes.
PersonaPlex is a real-time, full-duplex speech-to-speech conversational model with persona control through text-based role prompts and audio-based voice conditioning.
Why Quantize?
The original model requires ~14GB VRAM (bf16), which exceeds consumer GPUs like the RTX 4070 (12GB). This 4-bit quantized version:
| Original (bf16) | Quantized (NF4) | |
|---|---|---|
| VRAM | ~14 GiB | ~9.6 GiB |
| GPU | A100 / H100 | RTX 4070+ (12GB) |
| torch.compile | Yes | Yes |
| CUDA graphs | Yes | Yes |
What's Quantized?
Only the main transformer's linear layers (attention projections + gating FFN) are quantized to 4-bit NF4. The following are kept in bf16 for quality:
- Mimi audio encoder/decoder
- Depformer (depth transformer)
- Embedding layers
- Output heads
Quick Start
Prerequisites
- Accept the PersonaPlex license (required for the base model assets)
- Set your HuggingFace token:
export HF_TOKEN=<YOUR_TOKEN>
Installation
git clone https://huggingface.co/brianmatzelle/personaplex-7b-v1-bnb-4bit
cd personaplex-7b-v1-bnb-4bit
pip install moshi/.
pip install bitsandbytes
Run (Live Server)
SSL_DIR=$(mktemp -d)
python -m moshi.server --ssl "$SSL_DIR" --quantize-4bit
Then open https://localhost:8998 in your browser.
Only one browser/client WebSocket session is accepted at a time; extra connections get close code 1013 / SESSION_BUSY.
For 8-bit inference (higher VRAM than 4-bit, typically better quality retention):
SSL_DIR=$(mktemp -d)
python -m moshi.server --ssl "$SSL_DIR" --quantize-8bit
Run (FastAPI Production Server β API-only by default)
PersonaPlex is delivered as an API-first product. The supported public integration surface is the documented WebSocket + REST API; the bundled browser demo UI is disabled by default and must be opted into explicitly for internal demos.
python -m moshi.fastapi_server --quantize-4bit --host 0.0.0.0 --port 8998
Public surface (only paths an integrator should depend on):
WS /api/chatβ binary voice/text streaming protocol (canonical buyer path)GET /β JSON discovery document describing this API surfaceGET /healthβ basic livenessGET /readyβ runtime + GPU readinessGET /configβ effective runtime config (includeswebui_enabled,public_surface)GET /metricsβ Prometheus metrics
Notes:
- Hitting
/api/chatin a browser tab is expected to fail; it is a WebSocket endpoint. GET /configreportswebui_enabled: falseandpublic_surface: "websocket+rest"in the default API-only mode.- The service runs one inference session at a time (exclusive slot). Each WebSocket must include
voice_prompt(when the server is configured with a voice library) and may includetext_prompt. Optional per-session overrides (query string):context(attention span, clamped to the server's--contextcap),audio_temperature(aliasaudio_temp),audio_top_k(aliasaudio_topk),text_temperature(aliastext_temp),text_top_k(aliastext_topk). Values are restored after the session ends so the next client gets a clean baseline. - Latency tuning: server logs a
session telemetryline every ~5 s withavg_step_ms,max_step_ms,rtf(real-time factor; <1.0 keeps up with audio),pcm_backlog_ms, and the LMlm_offset/KV-cache fill. Ifrtfrises above 1.0 you'll also seesession falling behind real-timewarnings and clients receive aBEHIND_REALTIMEJSON diagnostic. Per-step attention cost grows linearly with the KV-cache fill untillm_offsetreaches the server'scontext(then stabilizes), so for latency-sensitive deployments lower the per-session?context=(or the server-wide--context).
Useful env overrides:
PERSONAPLEX_PORTPERSONAPLEX_HOSTPERSONAPLEX_SESSION_IDLE_TIMEOUT_SECONDSPERSONAPLEX_SESSION_ACQUIRE_WAIT_SECONDS(default60) β how long a second WebSocket may wait for the exclusive session slot. After this wait, close code1013/SESSION_BUSY.PERSONAPLEX_ENABLE_WEBUIβ set to1to mount the bundled browser demo UI at/(off by default).PERSONAPLEX_WEBUI_DISTβ optional path to the static UI folder used when the demo UI is enabled (defaults to./webui/dist).PERSONAPLEX_UVICORN_FULL_WS_URLβ set to1to restore full WebSocket URLs in uvicorn access/error logs (default is shortened; JSON logs still show applied sampling).PERSONAPLEX_LOG_LEVEL
Test as a client (no UI required)
Buyers integrate against the WebSocket protocol, not against any HTML/JS the server might ship. The repo includes a thin reference client that proves a deployment is healthy end-to-end without ever opening a browser.
- Quick REST sanity:
curl http://<host>:<port>/ # discovery JSON (API-only mode)
curl http://<host>:<port>/health
curl http://<host>:<port>/ready
curl http://<host>:<port>/config
- Stateless pre-flight check (no session slot consumed):
curl 'http://<host>:<port>/api/probe?voice_prompt=NATF2.pt' \
-H 'x-internal-auth: <token-if-configured>'
# Returns 200 + {ready: true, ...} when auth, voice_prompt, and the
# inference slot are all good; otherwise 409 with which field failed.
- WebSocket handshake smoke test (no audio):
# Either of these confirms the WS upgrade succeeds and (when configured)
# auth is correct. They will exit because they don't speak the binary
# protocol -- that is expected.
websocat -k 'ws://<host>:<port>/api/chat?voice_prompt=NATF2.pt'
# or
wscat -c 'ws://<host>:<port>/api/chat?voice_prompt=NATF2.pt'
Postman/wscat note: the WebSocket is a binary audio protocol. Any bad frame content -- text frame, empty binary frame, or unknown binary
kindbyte -- is reported as a non-fatal JSON warning text frame back from the server, and the session stays open so you can correct the next frame and continue. Example after sending the textHello:{"warning": "TEXT_FRAMES_NOT_SUPPORTED", "fatal": false, "message": "WS /api/chat is a binary audio protocol. Send 0x01 + Opus(24 kHz mono) frames. This text frame was ignored; the session is still open.", "hint": "See GET / for the frame protocol; use moshi.api_client as a reference."}The session is only closed for real problems: auth failure (
1008), capacity/setup (1013), session timeout or internal error (1011). Postman is therefore useful for connectivity, auth, and protocol-validation experiments; it cannot drive a real voice session because it cannot generate Opus audio.
- Postman-friendly audio in / WAV out (
POST /api/voice/onceβ no Opus coding on your side):
Formats
- WAV / FLAC / AIFF (anything
libsndfilecan read): send as raw body;Content-Typecan beaudio/wavor evenapplication/octet-stream. - WebM / MP4 (the options Postman shows, e.g.
audio/webm,video/webm,audio/mp4,video/mp4): set Headers β Content-Type to match the file you attach. The server decodes these with ffmpeg (theffmpegbinary must be on the serverPATH, e.g.apt install ffmpegon Ubuntu).
# curl: send a WAV, get a WAV back
curl -X POST 'http://<host>:<port>/api/voice/once?voice_prompt=NATF2.pt' \
-H 'Content-Type: audio/wav' \
-H 'x-internal-auth: <token-if-configured>' \
--data-binary @assets/test/input_assistant.wav \
-D /tmp/headers.txt \
-o agent_reply.wav
# curl: send a WebM recording (after installing ffmpeg on the server)
curl -X POST 'http://<host>:<port>/api/voice/once?voice_prompt=NATF2.pt' \
-H 'Content-Type: audio/webm' \
--data-binary @recording.webm \
-D /tmp/headers.txt \
-o agent_reply.wav
# Transcript JSON is URL-encoded in the x-transcript response header
python -c "from urllib.parse import unquote; import re; print(unquote([l for l in open('/tmp/headers.txt') if l.lower().startswith('x-transcript:')][0].split(': ', 1)[1].strip()))"
In Postman: POST to the same URL, Body β binary, pick your file. Set Headers β Content-Type to audio/webm, video/webm, audio/mp4, or video/mp4 when you are not sending WAV. The response body is always a WAV (use Save Response β Save to file and play locally); x-transcript holds the URL-encoded JSON transcript.
If WebM/MP4 fails with 503 / FFMPEG_NOT_AVAILABLE, install ffmpeg on the inference host. If soundfile cannot open the file and ffmpeg is missing, you get 400 / INVALID_AUDIO_BODY.
This endpoint is synchronous and shares the single-session lock with
/api/chat, so it is meant for testing and batch evaluation, not for
real-time concurrent traffic.
- Full voice session via the bundled reference client (
moshi.api_client):
pip install moshi/. # picks up websockets + soundfile
# Headless: drive a WAV file in, capture agent audio + transcript out
python -m moshi.api_client \
--url ws://<host>:<port>/api/chat \
--voice-prompt NATF2.pt \
--text-prompt "You are a friendly assistant." \
--input-wav assets/test/input_assistant.wav \
--output-wav agent_reply.wav
# Or live mic + speaker (requires a working audio device)
python -m moshi.api_client \
--url ws://<host>:<port>/api/chat \
--voice-prompt NATF2.pt \
--mic --speaker --duration 30
Streaming text tokens print to stdout as the agent talks; received audio
goes to --output-wav and/or the speaker. Pass --internal-auth-token β¦
when the deployment is configured with a shared secret, and --tenant-id β¦ to populate the x-tenant-id header. After install, the same client is
also available as personaplex-client β¦.
The WebSocket binary protocol is summarized in moshi/moshi/api_client.py's
module docstring; that file is the single source of truth for what
integrators have to implement.
Standalone desktop client folder (
clients/womens-helpline-desktop/) β a self-contained Women's Healthcare Helpline demo app you can copy to any laptop without cloning this repo. The inference stack stays remote on the server; only this folder needs to live on the tester's machine.What's in the folder:
app.pyβ single-file Tkinter desktop client (Start Call / End Call button, status pill, mm:ss timer, live transcript).requirements.txtβ 4 runtime deps (numpy,sphn,websockets,sounddevice). Notorch,fastapi, ormoshi.README.mdβ 1-page setup + run instructions for the tester.
On the tester's laptop:
# copy or download just clients/womens-helpline-desktop/ cd womens-helpline-desktop python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate pip install -r requirements.txt # On Linux only: sudo apt install python3-tk python app.py \ --url ws://<host>:<port>/api/chat \ --voice-prompt NATF2.pt # optional: --internal-auth-token <token>The persona prompt defaults to a women's-healthcare-helpline tone (warm, non-judgmental, never diagnoses, defers to qualified doctors). The tester can override it directly in the UI before clicking Start Call, or via
--text-prompt.
Productization: how to ship this to your end customers
Short answer: don't ship the raw WebSocket URL to end customers as the primary integration surface. It works for back-end-to-back-end usage and for advanced integrators, but a "click a button to talk to the agent" experience needs a thinner abstraction in front of it. The intended architecture (mirroring what Vapi, Retell, ElevenLabs, Daily, Twilio do) has four layers, in order of who owns each one:
| Layer | Who owns it | What it does |
|---|---|---|
| Inference + WS API | You (this repo) | POST /api/voice/once, WS /api/chat, /api/probe -- the binary contract integrators can program against. |
| Control plane (token mint) | You | An auth/billing service buyers' backends call (POST /v1/sessions) to mint short-lived per-session tokens. |
| Client SDK (npm / pip / Swift) | You | Wraps WebAudio + Opus encode/decode + WS; exposes a connect(token) API with start()/stop() events. |
| Embeddable widget | You (or buyer) | One-line <script>/<iframe> drop-in for non-engineering buyers. |
| Customer app | Customer | Calls SDK or drops in widget. Never sees the raw WS contract. |
For the AWS Marketplace launch in the project plan, the practical sequencing is:
- Now (done in this repo) β harden the API surface (WS + REST + probe + diagnostics).
- Next (Phase 2/3 of the plan) β build the control plane: per-tenant
subscription store,
ResolveCustomer/GetEntitlements, short-lived session token issuance,MeterUsagefor billing. The buyer's backend exchanges their AWS Marketplace customer identifier for a session token, then hands that token to the buyer's frontend. The frontend uses the token inx-internal-auth(or a successor header you choose for the public API) when opening/api/chat. - Then β publish a JS SDK (
@personaplex/web):await PersonaPlex.connect({ token })handles mic capture, Opus encoding, WS lifecycle, audio playback, and exposes events (onTranscript,onAudio,onConnected,onDisconnected). - Optional β a one-line embed (
<script src="https://cdn.personaplex.ai/embed.js" data-token-endpoint="...">) that drops a "Talk to agent" button into any page.
Why not skip straight to "give the customer the WebSocket URL":
- Browsers can't easily produce Opus-encoded audio; an SDK is required.
- You can't gate per-tenant usage, rotate tokens, or rate-limit at the WS level cleanly -- that belongs in a session-token mint endpoint.
- Marketplace billing requires
MeterUsagecalls that depend on your control plane knowing when a session started/ended -- the SDK is the natural place to emit those signals.
The current repo is the right foundation. The pieces above are explicit follow-on work and live in Phases 2-4 of the project plan.
Production tuning for real-time voice latency
Real-time voice has a tight latency budget (~150β300 ms is the threshold for "feels live"). Server inference time is only one component of that budget; the rest is dominated by network RTT and audio buffering on the client. Here is the playbook we use, in order of impact.
- Co-locate the server with your users. Cross-continent RTT alone can be 200β400 ms. Run the inference deployment in the same AWS region as the buyer's customers. For a multi-region rollout, front the API with GeoDNS or Route 53 latency-based routing.
- Tune
--contextto the call length you actually serve. Per-step attention cost grows linearly withmin(state.offset, context)until the KV cache fills. With--context 4000at 12.5 fps the cache fills in ~320 s and worst-case attention runs over 4000 tokens. For a typical 30β60 s helpline call, set the per-session?context=to 500β1500 β the cache fills in 40β120 s and the steady-state per-step GPU time is much lower. Server logs showavg_step_msandrtfper session so you can pick the right value empirically. - Watch the per-session telemetry. Every ~5 s the server emits a
session telemetryline withavg_step_ms,max_step_ms,rtf(real-time factor; <1 is healthy),pcm_backlog_ms, andlm_offset. Whenrtfis consistently <0.5 the server is the easy part of the pipeline. When it climbs above ~0.85 you have ~15% headroom left and should lowercontext. The metrics are also exposed at/metricsaspersonaplex_lm_step_secondsandpersonaplex_pcm_backlog_ms. - TCP_NODELAY is on by default for
/api/chatconnections (the server sets it afteraccept()). Don't put a Layer-7 proxy in front of the WebSocket that re-buffers small frames. nginx is fine if you setproxy_buffering offandproxy_request_buffering offon the WebSocket location (the bundleddeploy/nginx/personaplex.confalready does). - Fix client-side audio buffering. The largest single win is forcing
the client's playback library to use a small buffer:
sounddevice.OutputStream(latency='low', blocksize=480)(Python; the bundled desktop demo already uses this)- WebAudio: build the SDK on top of
AudioWorkletwith a 128- or 256-sample worklet processor, not a<audio>tag (which buffers hundreds of ms). - Native iOS/Android: use the low-latency audio APIs (
AVAudioEngine,AAudio) with the smallest supported buffer size; do not useAVAudioPlayerorMediaPlayer.
- Cap concurrency at the model. This server runs one session at a
time per process (the exclusive session gate). For more concurrent
callers, run multiple processes (one per GPU) behind a router that
round-robins by
tenant_idor by pure least-loaded; do not stack sessions on a single GPU process β KV-cache memory and step latency both scale with concurrent sessions and you'll lose real-time guarantees. - Add a health-aware load balancer.
GET /readyreturns 503 until warmup completes; point your LB's health check at it so a freshly booted instance doesn't take traffic before the GPU is hot. - Pin your CUDA + PyTorch versions and use CUDA graphs (already
enabled via
CUDAGraphedinLMGen). Driver/torch upgrades can shift per-step latency by 10β30 ms; treat them as deployment events with pre/post-deployment latency benchmarks. - Observability you should add for production.
- Structured logs (this repo emits JSON already) shipped to your log
pipeline. Always include
request_id,session_id,tenant_id. /metricsscraped by Prometheus + Grafana with alerts onpersonaplex_pcm_backlog_msp99 > 500 ms,personaplex_lm_step_secondsp99 > 70 ms, andpersonaplex_active_sessionsstuck at 1 for >5 min (deadlocked session).- Per-call recording of the
BEHIND_REALTIMEdiagnostic count.
- Structured logs (this repo emits JSON already) shipped to your log
pipeline. Always include
- At the contract layer, version the protocol. Today the WebSocket
speaks one binary protocol; the moment a buyer integrates against
/api/chat, you cannot change frame shapes without a coordinated cutover. Add a?protocol_version=query parameter and a JSON diagnostic on connect that includes the negotiated version, so future breaking changes can be additive.
If you ever see rtf < 0.5 in telemetry but the user still reports
"agent feels slow", the bottleneck is not the server β it's RTT, the
client's playback buffer, or end-pointing (the model's decision about
when the user has finished speaking). Diagnose in that order.
Optional: enable the bundled demo UI (internal use only)
The repo ships webui/dist (PersonaPlex browser UI). It is not the
buyer-facing integration path β buyers integrate against the documented
WebSocket protocol. To run the demo locally for evaluation:
PERSONAPLEX_ENABLE_WEBUI=1 \
python -m moshi.fastapi_server --quantize-4bit --host 0.0.0.0 --port 8998
# or equivalently:
python -m moshi.fastapi_server --quantize-4bit --enable-webui
When enabled, the server serves webui/dist at /. If the bundle is missing
locally, it is downloaded from Hugging Face (dist.tgz). For production we
recommend hosting any demo UI on a separate hostname and never mixing it
with the paid API surface.
Deploy Behind Nginx (TLS + WebSocket)
For public usage, terminate TLS at Nginx and proxy to the FastAPI process.
- Start the inference server on loopback:
python -m moshi.fastapi_server \
--quantize-4bit \
--host 127.0.0.1 \
--port 8004 \
--moshi-weight /home/user/personaplex-7b-v1-bnb-4bit/model_bnb_4bit_finetuned.pt \
--internal-auth-token "change-me"
- Install Nginx config:
sudo cp deploy/nginx/personaplex.conf /etc/nginx/sites-available/personaplex
sudo ln -sf /etc/nginx/sites-available/personaplex /etc/nginx/sites-enabled/personaplex
sudo nginx -t
sudo systemctl reload nginx
- Edit
deploy/nginx/personaplex.confvalues before reload:
- replace
api.example.comwith your real domain - replace cert paths with your real certificate files
- ensure
upstream personaplex_fastapipoints at your FastAPI listen address (default127.0.0.1:8004)
- Customer-facing URLs (API-only, no UI):
- Streaming API (WebSocket):
wss://api.example.com/api/chat?text_prompt=...&voice_prompt=... - API discovery:
https://api.example.com/ - REST health/readiness/config:
https://api.example.com/health,/ready,/config - Prometheus metrics (operator-only):
https://api.example.com/metrics
The shipped deploy/nginx/personaplex.conf proxies only the API paths and
returns 404 for everything else. The bundled browser demo UI is intentionally
not exposed via this config; if you want to run a demo, host it on a separate
hostname/port and start the backend with --enable-webui (see above).
Notes:
GET /api/chatin a browser tab is expected to fail; it is a WebSocket endpoint.- If using an IP instead of a domain, browser TLS and audio APIs can behave inconsistently. Use a domain + valid cert for production.
Run (Offline Evaluation)
python -m moshi.offline \
--voice-prompt "NATF2.pt" \
--input-wav "assets/test/input_assistant.wav" \
--seed 42424242 \
--output-wav "output.wav" \
--output-text "output.json" \
--quantize-4bit
8-bit offline variant:
python -m moshi.offline \
--voice-prompt "NATF2.pt" \
--input-wav "assets/test/input_assistant.wav" \
--seed 42424242 \
--output-wav "output.wav" \
--output-text "output.json" \
--quantize-8bit
Using Pre-Quantized Weights
This repo includes pre-quantized weights (model_bnb_4bit.pt) so you don't need the
full 16.7GB download. To use them, pass --moshi-weight model_bnb_4bit.pt along with
--quantize-4bit. The loader auto-detects the pre-quantized format and skips re-quantization.
Changes from Base Model
This repo includes a modified moshi/ package with:
--quantize-4bitflag for on-the-fly 4-bit NF4 quantization via bitsandbytes--quantize-8bitflag for on-the-fly 8-bit int8 quantization via bitsandbytes- Pre-quantized checkpoint loading (auto-detected, no re-quantization needed)
--cpu-offloadfixes for consumer GPU compatibility- Attention
in_projrefactored as a propernn.Modulefor quantization support - Gating forward path updated to route through quantized modules
Citation
@misc{roy2026personaplexvoicerolecontrol,
title={PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models},
author={Rajarshi Roy and Jonathan Raiman and Sang-gil Lee and Teodor-Dumitru Ene and Robert Kirby and Sungwon Kim and Jaehyeon Kim and Bryan Catanzaro},
year={2026},
eprint={2602.06053},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.06053},
}
License
Code is MIT licensed. Model weights are under the NVIDIA Open Model License.