Download mixed/README.md from soniqo/PersonaPlex-7B-ONNX: direct link, hf CLI and curl.
- Browser
- Download file 3.87 kB
-
https://huggingface.co/soniqo/PersonaPlex-7B-ONNX/resolve/55a31aba70f456e532a988b3af0db1148b3e0220/mixed/README.md
- Command line
-
hf download hf://soniqo/PersonaPlex-7B-ONNX@55a31aba70f456e532a988b3af0db1148b3e0220/mixed/README.md
-
curl -L -o README.md https://huggingface.co/soniqo/PersonaPlex-7B-ONNX/resolve/55a31aba70f456e532a988b3af0db1148b3e0220/mixed/README.md
license: other
license_name: nvidia-open-model-license
license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE
language:
- en
library_name: onnxruntime
tags:
- personaplex
- speech-to-speech
- full-duplex
- voice-agent
- onnx
- moshi-architecture
base_model: nvidia/personaplex-7b-v1
PersonaPlex 7B ONNX -- mixed variant
Pareto winner on quality + VRAM. INT8-dynamic temporal + FP16 depformer + FP32 mimi.
This is one of four ONNX-quantized bundles of NVIDIA's PersonaPlex 7B -- a full-duplex speech-to-speech model on Kyutai's Moshi architecture. The full collection is at soniqo/PersonaPlex-7B-ONNX.
Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames)
| This bundle | |
|---|---|
| Disk | ~11 GB |
| Host RAM | 7.9 GB |
| VRAM | 6.6 GB |
| RTF (12.5 Hz) | 3.5x (1.0 = realtime) |
| Quality | cos 0.990448 vs FP32 reference |
Best quality at low VRAM. RTF is the cost of INT8 dynamic Q/DQ.
All four variants
| Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for |
|---|---|---|---|---|---|---|
fp16 |
~17 GB | 1.5 GB | 18.3 GB | 5.3x | 1.0000 | Near-perfect quality |
mixed |
~11 GB | 7.9 GB | 6.6 GB | 3.5x | 0.9904 | Best quality at low VRAM ← (this) |
int8-nb-dep_gint8 ⭐ |
~9.4 GB | 1.4 GB | 12.1 GB | 1.12x | 0.9977 | Best balance: near realtime RTF, low host RAM, excellent quality |
int4-nb-dep_gint8 |
~7.6 GB | 1.4 GB | 9.6 GB | 1.12x | 0.8774 | Smallest viable bundle |
The int8-nb-dep_gint8 variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998).
The mixed variant wins on quality + VRAM combined but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint.
How to use
This bundle is consumed by the speech-core C++ inference runtime:
# Download (from speech-core repo root)
PERSONAPLEX_VARIANT=mixed scripts/download_personaplex_onnx.sh
# Run
build/Release/run_personaplex scripts/personaplex-mixed 50 audio.wav VARF2
Or load directly with ONNX Runtime in any language:
import onnxruntime as ort
sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"])
# inputs: text_token [1,1] int64, audio_tokens [1,16] int64,
# past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float
# outputs: hidden [1,1,4096] float, new_k_all, new_v_all
Files in this variant
| File | Purpose |
|---|---|
mimi_encoder.onnx(+.data) |
24 kHz PCM -> 16 audio codebooks @ 12.5 Hz |
mimi_decoder.onnx(+.data) |
16 audio codebooks @ 12.5 Hz -> 24 kHz PCM |
temporal_step.onnx(+.data) |
One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O |
depformer_step.onnx(+.data) |
One inner step of 6-layer depformer, 16 codebook steps per frame |
tokenizer_spm_32k_3.model |
SentencePiece text tokenizer |
voices/<name>.bin |
18 voice prompts (NATF*, NATM*, VARF*, VARM*) |
system_prompts.bin |
Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts |
config.json |
Architecture + precision + measured-metrics metadata |
Related
- soniqo/speech-core -- C++ runtime that consumes this bundle, with the
OnnxPersonaPlexwrapper and CUDA EP routing - soniqo.audio -- the project site
License
Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License).