File size: 3,848 Bytes
55a31ab | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 | ---
license: other
license_name: nvidia-open-model-license
license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE
language:
- en
library_name: onnxruntime
tags:
- personaplex
- speech-to-speech
- full-duplex
- voice-agent
- onnx
- moshi-architecture
base_model: nvidia/personaplex-7b-v1
---
# PersonaPlex 7B ONNX -- `fp16` variant
Maximum quality, maximum VRAM. FP16 temporal + FP16 depformer + FP32 mimi.
This is one of four ONNX-quantized bundles of NVIDIA's
[PersonaPlex 7B](https://huggingface.co/nvidia/personaplex-7b-v1) -- a
full-duplex speech-to-speech model on Kyutai's
[Moshi](https://github.com/kyutai-labs/moshi) architecture. The full
collection is at
[soniqo/PersonaPlex-7B-ONNX](https://huggingface.co/soniqo/PersonaPlex-7B-ONNX).
## Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames)
| | This bundle |
|---|---|
| **Disk** | ~17 GB |
| **Host RAM** | 1.5 GB |
| **VRAM** | 18.3 GB |
| **RTF (12.5 Hz)** | 5.3x (1.0 = realtime) |
| **Quality** | cos 0.999999 vs FP32 reference |
Near-perfect quality. Use when VRAM is plentiful.
## All four variants
| Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for |
|---|---|---|---|---|---|---|
| `fp16` | ~17 GB | 1.5 GB | 18.3 GB | 5.3x | 1.0000 | Near-perfect quality ← **(this)** |
| `mixed` | ~11 GB | 7.9 GB | 6.6 GB | 3.5x | 0.9904 | Best quality at low VRAM |
| `int8-nb-dep_gint8` ⭐ | ~9.4 GB | 1.4 GB | 12.1 GB | 1.12x | 0.9977 | Best balance: near realtime RTF, low host RAM, excellent quality |
| `int4-nb-dep_gint8` | ~7.6 GB | 1.4 GB | 9.6 GB | 1.12x | 0.8774 | Smallest viable bundle |
The `int8-nb-dep_gint8` variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998).
The `mixed` variant wins on **quality + VRAM combined** but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint.
## How to use
This bundle is consumed by the [speech-core](https://github.com/soniqo/speech-core) C++ inference runtime:
```bash
# Download (from speech-core repo root)
PERSONAPLEX_VARIANT=fp16 scripts/download_personaplex_onnx.sh
# Run
build/Release/run_personaplex scripts/personaplex-fp16 50 audio.wav VARF2
```
Or load directly with ONNX Runtime in any language:
```python
import onnxruntime as ort
sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"])
# inputs: text_token [1,1] int64, audio_tokens [1,16] int64,
# past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float
# outputs: hidden [1,1,4096] float, new_k_all, new_v_all
```
## Files in this variant
| File | Purpose |
|---|---|
| `mimi_encoder.onnx`(+`.data`) | 24 kHz PCM -> 16 audio codebooks @ 12.5 Hz |
| `mimi_decoder.onnx`(+`.data`) | 16 audio codebooks @ 12.5 Hz -> 24 kHz PCM |
| `temporal_step.onnx`(+`.data`) | One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O |
| `depformer_step.onnx`(+`.data`) | One inner step of 6-layer depformer, 16 codebook steps per frame |
| `tokenizer_spm_32k_3.model` | SentencePiece text tokenizer |
| `voices/<name>.bin` | 18 voice prompts (NATF*, NATM*, VARF*, VARM*) |
| `system_prompts.bin` | Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts |
| `config.json` | Architecture + precision + measured-metrics metadata |
## Related
- [soniqo/speech-core](https://github.com/soniqo/speech-core) -- C++ runtime that consumes this bundle, with the `OnnxPersonaPlex` wrapper and CUDA EP routing
- [soniqo.audio](https://soniqo.audio) -- the project site
## License
Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License).
|