File size: 3,874 Bytes
e86dfd1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
---

license: other
license_name: nvidia-open-model-license
license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE
language:
- en
library_name: onnxruntime
tags:
- personaplex
- speech-to-speech
- full-duplex
- voice-agent
- onnx
- moshi-architecture
base_model: nvidia/personaplex-7b-v1
---


# PersonaPlex 7B ONNX -- `mixed` variant

Pareto winner on quality + VRAM. INT8-dynamic temporal + FP16 depformer + FP32 mimi.

This is one of four ONNX-quantized bundles of NVIDIA's
[PersonaPlex 7B](https://huggingface.co/nvidia/personaplex-7b-v1) -- a
full-duplex speech-to-speech model on Kyutai's
[Moshi](https://github.com/kyutai-labs/moshi) architecture. The full
collection is at
[soniqo/PersonaPlex-7B-ONNX](https://huggingface.co/soniqo/PersonaPlex-7B-ONNX).

## Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames)

| | This bundle |
|---|---|
| **Disk** | ~11 GB |
| **Host RAM** | 7.9 GB |
| **VRAM** | 6.6 GB |
| **RTF (12.5 Hz)** | 3.5x (1.0 = realtime) |
| **Quality** | cos 0.990448 vs FP32 reference |

Best quality at low VRAM. RTF is the cost of INT8 dynamic Q/DQ.

## All four variants

| Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for |
|---|---|---|---|---|---|---|
| `fp16` | ~17 GB | 1.5 GB | 18.3 GB | 5.3x | 1.0000 | Near-perfect quality |
| `mixed` | ~11 GB | 7.9 GB | 6.6 GB | 3.5x | 0.9904 | Best quality at low VRAM ← **(this)** |
| `int8-nb-dep_gint8` ⭐ | ~9.4 GB | 1.4 GB | 12.1 GB | 1.12x | 0.9977 | Best balance: near realtime RTF, low host RAM, excellent quality |
| `int4-nb-dep_gint8` | ~7.6 GB | 1.4 GB | 9.6 GB | 1.12x | 0.8774 | Smallest viable bundle |

The `int8-nb-dep_gint8` variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998).

The `mixed` variant wins on **quality + VRAM combined** but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint.

## How to use

This bundle is consumed by the [speech-core](https://github.com/soniqo/speech-core) C++ inference runtime:

```bash

# Download (from speech-core repo root)

PERSONAPLEX_VARIANT=mixed scripts/download_personaplex_onnx.sh



# Run

build/Release/run_personaplex scripts/personaplex-mixed 50 audio.wav VARF2

```

Or load directly with ONNX Runtime in any language:

```python

import onnxruntime as ort

sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"])

# inputs: text_token [1,1] int64, audio_tokens [1,16] int64,

#         past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float

# outputs: hidden [1,1,4096] float, new_k_all, new_v_all

```

## Files in this variant

| File | Purpose |
|---|---|
| `mimi_encoder.onnx`(+`.data`) | 24 kHz PCM -> 16 audio codebooks @ 12.5 Hz |
| `mimi_decoder.onnx`(+`.data`) | 16 audio codebooks @ 12.5 Hz -> 24 kHz PCM |
| `temporal_step.onnx`(+`.data`) | One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O |
| `depformer_step.onnx`(+`.data`) | One inner step of 6-layer depformer, 16 codebook steps per frame |
| `tokenizer_spm_32k_3.model` | SentencePiece text tokenizer |
| `voices/<name>.bin` | 18 voice prompts (NATF*, NATM*, VARF*, VARM*) |
| `system_prompts.bin` | Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts |
| `config.json` | Architecture + precision + measured-metrics metadata |

## Related

- [soniqo/speech-core](https://github.com/soniqo/speech-core) -- C++ runtime that consumes this bundle, with the `OnnxPersonaPlex` wrapper and CUDA EP routing
- [soniqo.audio](https://soniqo.audio) -- the project site

## License

Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License).