aufklarer commited on
Commit
f36d5d6
Β·
verified Β·
1 Parent(s): 55a31ab

Add top-level model card listing all four variants

Browse files
Files changed (1) hide show
  1. README.md +128 -0
README.md ADDED
@@ -0,0 +1,128 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: nvidia-open-model-license
4
+ license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE
5
+ language:
6
+ - en
7
+ library_name: onnxruntime
8
+ tags:
9
+ - personaplex
10
+ - speech-to-speech
11
+ - full-duplex
12
+ - voice-agent
13
+ - onnx
14
+ - moshi-architecture
15
+ base_model: nvidia/personaplex-7b-v1
16
+ ---
17
+
18
+ # PersonaPlex 7B ONNX
19
+
20
+ ONNX-quantized bundles of NVIDIA's [PersonaPlex 7B](https://huggingface.co/nvidia/personaplex-7b-v1) β€” a full-duplex speech-to-speech model on Kyutai's [Moshi](https://github.com/kyutai-labs/moshi) architecture. Listens and speaks simultaneously at 12.5 Hz, conditioned on a voice preset and a text system prompt.
21
+
22
+ This repository ships **four production-ready bundle variants** spanning the disk Γ— host RAM Γ— VRAM Γ— RTF Γ— quality trade-off space. Pick one based on your target hardware and quality bar.
23
+
24
+ ## Variants at a glance (measured on RTX 5090, 50 frames, VARF2 voice, "helpful" prompt)
25
+
26
+ | Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for |
27
+ |---|---|---|---|---|---|---|
28
+ | [**`int8-nb-dep_gint8`**](./int8-nb-dep_gint8/) ⭐ | 9.4 GB | **1.4 GB** | 12.1 GB | **1.12Γ—** | 0.998 | **Recommended ship default** β€” best RTF + low host RAM + excellent quality |
29
+ | [`mixed`](./mixed/) | 11 GB | 7.9 GB | **6.6 GB** | 3.5Γ— | 0.990 | **Quality + VRAM Pareto winner** β€” lowest VRAM + best topical output ("We're concerned about it.") |
30
+ | [`int4-nb-dep_gint8`](./int4-nb-dep_gint8/) | **7.6 GB** | 1.4 GB | **9.6 GB** | 1.12Γ— | 0.877 | Smallest disk + lowest VRAM combo. Coherent but visibly degraded |
31
+ | [`fp16`](./fp16/) | 17 GB | 1.5 GB | 18.3 GB | 5.3Γ— | 0.9999 | Near-perfect quality, max VRAM |
32
+
33
+ RTF (real-time factor) is per-frame latency / frame interval at 12.5 Hz β€” **1.0Γ— = exactly realtime**, < 1.0Γ— = faster than realtime. The `int8-nb-dep_gint8` and `int4-nb-dep_gint8` variants run at ~1.12Γ— (near-realtime streaming).
34
+
35
+ ## Which variant to pick
36
+
37
+ - **Most use cases β†’ `int8-nb-dep_gint8`**: best balance of RTF, host RAM, and quality. If your GPU has β‰₯16 GB VRAM, this is what you want.
38
+ - **Limited VRAM (≀8 GB GPU) β†’ `mixed`**: only 6.6 GB VRAM. Costs host RAM and RTF, but quality is still excellent (cos 0.990) with the best topical responses on our benchmark.
39
+ - **Disk-constrained β†’ `int4-nb-dep_gint8`**: 7.6 GB on disk and only 9.6 GB VRAM. Accept some quality drift (cos 0.877 β€” coherent English but less precise).
40
+ - **Maximum quality regardless of cost β†’ `fp16`**: cos 0.9999, indistinguishable from the FP32 reference.
41
+
42
+ ## Architecture
43
+
44
+ ```
45
+ [User audio 24 kHz PCM]
46
+ ↓
47
+ [Mimi encoder: SEANet + 8L transformer + RVQ] β†’ 16 codebooks @ 12.5 Hz
48
+ ↓
49
+ [Temporal transformer: 32L, dim=4096, 7B params, RoPE, RMSNorm, SwiGLU]
50
+ ↓
51
+ [Depformer: 6L, dim=1024, MultiLinear Γ— 16 codebook steps] β†’ 16 agent audio tokens
52
+ ↓
53
+ [Mimi decoder] β†’ 24 kHz agent audio PCM
54
+ ```
55
+
56
+ Each variant ships four ONNX graphs:
57
+
58
+ | File | Purpose |
59
+ |---|---|
60
+ | `mimi_encoder.onnx`(+`.data`) | 24 kHz PCM β†’ 16 audio codebooks @ 12.5 Hz |
61
+ | `mimi_decoder.onnx`(+`.data`) | 16 audio codebooks @ 12.5 Hz β†’ 24 kHz PCM |
62
+ | `temporal_step.onnx`(+`.data`) | One frame of the 32-layer 7B temporal transformer, explicit KV-cache I/O |
63
+ | `depformer_step.onnx`(+`.data`) | One inner step of the 6-layer depformer, 16 codebook steps per frame |
64
+
65
+ Plus per-variant auxiliary files:
66
+
67
+ | File | Purpose |
68
+ |---|---|
69
+ | `tokenizer_spm_32k_3.model` | SentencePiece text tokenizer |
70
+ | `voices/<name>.bin` | 18 voice prompts (NATF0-3, NATM0-3, VARF0-4, VARM0-4) |
71
+ | `system_prompts.bin` | Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts |
72
+ | `config.json` | Architecture + precision + measured metrics |
73
+
74
+ ## How to use
75
+
76
+ ### Via the C++ runtime ([speech-core](https://github.com/soniqo/speech-core))
77
+
78
+ ```bash
79
+ # Download the recommended variant
80
+ PERSONAPLEX_VARIANT=int8-nb-dep_gint8 scripts/download_personaplex_onnx.sh
81
+
82
+ # Run end-to-end
83
+ build/Release/run_personaplex scripts/personaplex-int8-nb-dep_gint8 50 \
84
+ tests/data/test_audio.wav VARF2
85
+ ```
86
+
87
+ ### Via ONNX Runtime in Python
88
+
89
+ ```python
90
+ import onnxruntime as ort
91
+
92
+ # Inputs: text_token [1,1] int64
93
+ # audio_tokens [1,16] int64
94
+ # past_k_all [32, 1, 32, T_past, 128] float (FP32 or FP16 depending on bundle)
95
+ # past_v_all (same shape)
96
+ # Outputs: hidden [1, 1, 4096]
97
+ # new_k_all [32, 1, 32, T_full, 128]
98
+ # new_v_all (same shape)
99
+ sess = ort.InferenceSession("temporal_step.onnx",
100
+ providers=["CUDAExecutionProvider"])
101
+ ```
102
+
103
+ For full-duplex generation, also call `depformer_step` 16 times per frame (one inner step per audio codebook) β€” see the [speech-core wrapper source](https://github.com/soniqo/speech-core/blob/main/src/models/personaplex/onnx_personaplex.cpp) for the complete loop.
104
+
105
+ ## How these bundles were produced
106
+
107
+ All four bundles export from the FP32 PyTorch reference via stages in [`convert_onnx.py`](https://github.com/soniqo/speech-models/blob/main/models/personaplex/export/convert_onnx.py):
108
+
109
+ | Variant | Temporal | Depformer | Notes |
110
+ |---|---|---|---|
111
+ | `fp16` | FP16 weights | FP16 weights | Standard `torch.onnx.export` at `--dtype float16` |
112
+ | `mixed` | INT8 dynamic via `quantize_dynamic` (per-channel, FP32 scales) | FP16 weights | The classic "INT8 mixed precision" recipe |
113
+ | `int8-nb-dep_gint8` | INT8 via `MatMulNBitsQuantizer(bits=8, block=128)` | Custom INT8 quantization of the depformer's 24 large 3D Gather-source weight tensors (4 GB depformer disk savings via [`quantize_depformer_gather.py`](https://github.com/soniqo/speech-models/blob/main/models/personaplex/export/quantize_depformer_gather.py)) | Best balance |
114
+ | `int4-nb-dep_gint8` | INT4 via `MatMulNBitsQuantizer(bits=4, block=32)` | Same custom INT8 depformer | Smallest |
115
+
116
+ Mimi codec is FP32 in all four variants (small enough not to matter).
117
+
118
+ ## Related
119
+
120
+ - [soniqo/speech-core](https://github.com/soniqo/speech-core) β€” C++ inference runtime with the `OnnxPersonaPlex` wrapper, CUDA EP routing, multi-turn KV cache, and 12 memory-tuning env knobs (`SPEECH_CORE_USE_ENV_ALLOCATORS`, etc.)
121
+ - [soniqo/speech-models](https://github.com/soniqo/speech-models) β€” model export pipeline including `convert_onnx.py`, `quantize_depformer_gather.py`, `bench_pytorch_cuda.py`, `compare_bundle_quality.py`
122
+ - [soniqo.audio](https://soniqo.audio) β€” the project site
123
+ - [nvidia/personaplex-7b-v1](https://huggingface.co/nvidia/personaplex-7b-v1) β€” upstream PyTorch reference
124
+ - [kyutai-labs/moshi](https://github.com/kyutai-labs/moshi) β€” base architecture
125
+
126
+ ## License
127
+
128
+ NVIDIA Open Model License (same as upstream). See [the LICENSE link](https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE).