aufklarer commited on
Commit
55a31ab
·
verified ·
1 Parent(s): e86dfd1

upload fp16 bundle

Browse files
.gitattributes CHANGED
@@ -45,3 +45,7 @@ mixed/depformer_step.onnx.data filter=lfs diff=lfs merge=lfs -text
45
  mixed/mimi_decoder.onnx.data filter=lfs diff=lfs merge=lfs -text
46
  mixed/mimi_encoder.onnx.data filter=lfs diff=lfs merge=lfs -text
47
  mixed/temporal_step.onnx.data filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
45
  mixed/mimi_decoder.onnx.data filter=lfs diff=lfs merge=lfs -text
46
  mixed/mimi_encoder.onnx.data filter=lfs diff=lfs merge=lfs -text
47
  mixed/temporal_step.onnx.data filter=lfs diff=lfs merge=lfs -text
48
+ fp16/depformer_step.onnx.data filter=lfs diff=lfs merge=lfs -text
49
+ fp16/mimi_decoder.onnx.data filter=lfs diff=lfs merge=lfs -text
50
+ fp16/mimi_encoder.onnx.data filter=lfs diff=lfs merge=lfs -text
51
+ fp16/temporal_step.onnx.data filter=lfs diff=lfs merge=lfs -text
fp16/README.md ADDED
@@ -0,0 +1,96 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: nvidia-open-model-license
4
+ license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE
5
+ language:
6
+ - en
7
+ library_name: onnxruntime
8
+ tags:
9
+ - personaplex
10
+ - speech-to-speech
11
+ - full-duplex
12
+ - voice-agent
13
+ - onnx
14
+ - moshi-architecture
15
+ base_model: nvidia/personaplex-7b-v1
16
+ ---
17
+
18
+ # PersonaPlex 7B ONNX -- `fp16` variant
19
+
20
+ Maximum quality, maximum VRAM. FP16 temporal + FP16 depformer + FP32 mimi.
21
+
22
+ This is one of four ONNX-quantized bundles of NVIDIA's
23
+ [PersonaPlex 7B](https://huggingface.co/nvidia/personaplex-7b-v1) -- a
24
+ full-duplex speech-to-speech model on Kyutai's
25
+ [Moshi](https://github.com/kyutai-labs/moshi) architecture. The full
26
+ collection is at
27
+ [soniqo/PersonaPlex-7B-ONNX](https://huggingface.co/soniqo/PersonaPlex-7B-ONNX).
28
+
29
+ ## Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames)
30
+
31
+ | | This bundle |
32
+ |---|---|
33
+ | **Disk** | ~17 GB |
34
+ | **Host RAM** | 1.5 GB |
35
+ | **VRAM** | 18.3 GB |
36
+ | **RTF (12.5 Hz)** | 5.3x (1.0 = realtime) |
37
+ | **Quality** | cos 0.999999 vs FP32 reference |
38
+
39
+ Near-perfect quality. Use when VRAM is plentiful.
40
+
41
+ ## All four variants
42
+
43
+ | Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for |
44
+ |---|---|---|---|---|---|---|
45
+ | `fp16` | ~17 GB | 1.5 GB | 18.3 GB | 5.3x | 1.0000 | Near-perfect quality ← **(this)** |
46
+ | `mixed` | ~11 GB | 7.9 GB | 6.6 GB | 3.5x | 0.9904 | Best quality at low VRAM |
47
+ | `int8-nb-dep_gint8` ⭐ | ~9.4 GB | 1.4 GB | 12.1 GB | 1.12x | 0.9977 | Best balance: near realtime RTF, low host RAM, excellent quality |
48
+ | `int4-nb-dep_gint8` | ~7.6 GB | 1.4 GB | 9.6 GB | 1.12x | 0.8774 | Smallest viable bundle |
49
+
50
+ The `int8-nb-dep_gint8` variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998).
51
+
52
+ The `mixed` variant wins on **quality + VRAM combined** but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint.
53
+
54
+ ## How to use
55
+
56
+ This bundle is consumed by the [speech-core](https://github.com/soniqo/speech-core) C++ inference runtime:
57
+
58
+ ```bash
59
+ # Download (from speech-core repo root)
60
+ PERSONAPLEX_VARIANT=fp16 scripts/download_personaplex_onnx.sh
61
+
62
+ # Run
63
+ build/Release/run_personaplex scripts/personaplex-fp16 50 audio.wav VARF2
64
+ ```
65
+
66
+ Or load directly with ONNX Runtime in any language:
67
+
68
+ ```python
69
+ import onnxruntime as ort
70
+ sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"])
71
+ # inputs: text_token [1,1] int64, audio_tokens [1,16] int64,
72
+ # past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float
73
+ # outputs: hidden [1,1,4096] float, new_k_all, new_v_all
74
+ ```
75
+
76
+ ## Files in this variant
77
+
78
+ | File | Purpose |
79
+ |---|---|
80
+ | `mimi_encoder.onnx`(+`.data`) | 24 kHz PCM -> 16 audio codebooks @ 12.5 Hz |
81
+ | `mimi_decoder.onnx`(+`.data`) | 16 audio codebooks @ 12.5 Hz -> 24 kHz PCM |
82
+ | `temporal_step.onnx`(+`.data`) | One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O |
83
+ | `depformer_step.onnx`(+`.data`) | One inner step of 6-layer depformer, 16 codebook steps per frame |
84
+ | `tokenizer_spm_32k_3.model` | SentencePiece text tokenizer |
85
+ | `voices/<name>.bin` | 18 voice prompts (NATF*, NATM*, VARF*, VARM*) |
86
+ | `system_prompts.bin` | Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts |
87
+ | `config.json` | Architecture + precision + measured-metrics metadata |
88
+
89
+ ## Related
90
+
91
+ - [soniqo/speech-core](https://github.com/soniqo/speech-core) -- C++ runtime that consumes this bundle, with the `OnnxPersonaPlex` wrapper and CUDA EP routing
92
+ - [soniqo.audio](https://soniqo.audio) -- the project site
93
+
94
+ ## License
95
+
96
+ Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License).
fp16/config.json ADDED
@@ -0,0 +1,125 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "personaplex",
3
+ "version": "personaplex-7b-v1",
4
+ "base_model": "nvidia/personaplex-7b-v1",
5
+ "bundle_variant": "fp16",
6
+ "bundle_size_gb": 17,
7
+ "precision": {
8
+ "temporal": "fp16",
9
+ "depformer": "fp16",
10
+ "mimi": "fp32"
11
+ },
12
+ "metrics": {
13
+ "host_ram_gb": 1.5,
14
+ "vram_gb": 18.3,
15
+ "rtf": 5.3,
16
+ "hidden_cos_vs_fp32": 0.999999,
17
+ "verdict": "Near-perfect quality. Use when VRAM is plentiful."
18
+ },
19
+ "description": "Maximum quality, maximum VRAM. FP16 temporal + FP16 depformer + FP32 mimi.",
20
+ "temporal": {
21
+ "dim": 4096,
22
+ "num_layers": 32,
23
+ "num_heads": 32,
24
+ "head_dim": 128,
25
+ "hidden_scale": 4.125,
26
+ "n_q": 16,
27
+ "card": 2048,
28
+ "text_card": 32000,
29
+ "context": 3000,
30
+ "max_period": 10000
31
+ },
32
+ "depformer": {
33
+ "dim": 1024,
34
+ "num_layers": 6,
35
+ "num_heads": 16,
36
+ "head_dim": 64,
37
+ "dim_feedforward": 4224,
38
+ "num_steps": 16,
39
+ "card": 2048,
40
+ "context": 8,
41
+ "weights_per_step": true,
42
+ "multi_linear": true,
43
+ "pos_emb": "none"
44
+ },
45
+ "mimi": {
46
+ "sample_rate": 24000,
47
+ "frame_rate": 12.5,
48
+ "num_codebooks_used": 8,
49
+ "num_codebooks_internal": 32,
50
+ "codebook_size": 2048,
51
+ "samples_per_frame": 1920
52
+ },
53
+ "delays": [
54
+ 0,
55
+ 0,
56
+ 1,
57
+ 1,
58
+ 1,
59
+ 1,
60
+ 1,
61
+ 1,
62
+ 1,
63
+ 0,
64
+ 1,
65
+ 1,
66
+ 1,
67
+ 1,
68
+ 1,
69
+ 1,
70
+ 1
71
+ ],
72
+ "sampling": {
73
+ "audio_temp": 0.8,
74
+ "audio_top_k": 250,
75
+ "text_temp": 0.7,
76
+ "text_top_k": 25
77
+ },
78
+ "voices": [
79
+ "NATF0",
80
+ "NATF1",
81
+ "NATF2",
82
+ "NATF3",
83
+ "NATM0",
84
+ "NATM1",
85
+ "NATM2",
86
+ "NATM3",
87
+ "VARF0",
88
+ "VARF1",
89
+ "VARF2",
90
+ "VARF3",
91
+ "VARF4",
92
+ "VARM0",
93
+ "VARM1",
94
+ "VARM2",
95
+ "VARM3",
96
+ "VARM4"
97
+ ],
98
+ "system_prompts": [
99
+ "helpful",
100
+ "expert",
101
+ "warm",
102
+ "direct"
103
+ ],
104
+ "voice_embedding": {
105
+ "frames": 50,
106
+ "dim": 4096,
107
+ "default_scale": 10.0
108
+ },
109
+ "prefill_layout": {
110
+ "voice_embedding_prefix_frames": 50,
111
+ "voice_cache_replay_frames": 4,
112
+ "silence_spacer_frames": 6
113
+ },
114
+ "graphs": {
115
+ "mimi_encoder": "mimi_encoder.onnx",
116
+ "mimi_decoder": "mimi_decoder.onnx",
117
+ "temporal_step": "temporal_step.onnx",
118
+ "depformer_step": "depformer_step.onnx"
119
+ },
120
+ "auxiliary": {
121
+ "tokenizer": "tokenizer_spm_32k_3.model",
122
+ "system_prompts": "system_prompts.bin",
123
+ "voices_dir": "voices/"
124
+ }
125
+ }
fp16/depformer_step.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f535fa2c0839e891b90fcb2bc1b4393b53d37fe575a16693ddc5d43272d5affe
3
+ size 63981
fp16/depformer_step.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3dcc9a16dc078e11942086cb39166bef92af311193714db10327ea3f7f8a3e7d
3
+ size 2796109824
fp16/mimi_decoder.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0fd6447a47893f7460a852a1178fe6a537301709bfc8d7cae5dc033bfb2bef96
3
+ size 67748881
fp16/mimi_decoder.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9822aa363c73101d58eaaa7ce1d33e5e953576be8823cd0c2fecf2dbacef0bb8
3
+ size 160713728
fp16/mimi_encoder.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f70c97e6aa4194403f4de76b3eaafd47913b942e1de4154f34b82aa9d34bb901
3
+ size 763072
fp16/mimi_encoder.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:16d14e7b15815b546de821d0adf75082254095e9db0c1e3278745aa1bbfdecf9
3
+ size 223886080
fp16/system_prompts.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:00ddeee95bd65c5dd6b9b75004546fd5ac595885acff49fb72ed0dd1134ce6ad
3
+ size 223
fp16/temporal_step.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b241a434243dd1cecb6661a2bc6d224fb556f813b74a942f57584638dd26df16
3
+ size 494439
fp16/temporal_step.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3115967ef99660c587fb40855b078cd60d7a44a36eb2f243413db148b5aaeb68
3
+ size 13685121024
fp16/tokenizer_spm_32k_3.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:78d4336533ddc26f9acf7250d7fb83492152196c6ea4212c841df76933f18d2d
3
+ size 552778
fp16/voices/NATF0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fd7e4c1e37b3dc5a5583c9ee0cd66a9c1e64a86cdfabcd0bed1d68bf93e8d61c
3
+ size 836152
fp16/voices/NATF1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9e885da22e8eb20b0cd2be99eb08cbc73fba78c73a7e0ecddb6359d94c7dea95
3
+ size 787000
fp16/voices/NATF2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b38778578029eaebb6d0a69b0b5dbb150b058fb7f7c7f9070fcf967d22a8ecbc
3
+ size 836152
fp16/voices/NATF3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:247d0ef94f463090bd7e2ba83d1aa475f2baf55b7d8cd09e95eb02eef22abcd5
3
+ size 836152
fp16/voices/NATM0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5a73ed98b4d7e0012b496d9859b16ba6a5eb9672f4f7bfad0253d25df5a61510
3
+ size 819768
fp16/voices/NATM1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c52210a631f20b378537e60ed1eeff3b9930b62faf0083249e4979d7c516102b
3
+ size 836152
fp16/voices/NATM2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ee7082e5fd5a11aab45109c7c20c4042bd87237cec413b83893bba2a1d765434
3
+ size 819768
fp16/voices/NATM3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:84b93a791b440862a9bea2c5c6a17b34fb2d595a10173a29195207112676258c
3
+ size 754232
fp16/voices/VARF0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6ae429e01a3e0d03313dc1749e6182a6360cb58909627673a3569e2669015f03
3
+ size 1114680
fp16/voices/VARF1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cebc82dcd2040250e767ba3d037a9e6b7cee5a5c07b66c939eb866197ff9aac0
3
+ size 819768
fp16/voices/VARF2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e10657ee1dde618a419c1d2ef10cb4677b02b7d5335cd90b82265a2ee37475d5
3
+ size 836152
fp16/voices/VARF3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:14edf219fdd46edf01a71b5f6c7ed06dfc0563bbaa920cb38002298d18aec816
3
+ size 1016376
fp16/voices/VARF4.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e8a095738ef56ebd3ea29c990c3f8a901b5748874d6538d3410a1a62e6a9e0ce
3
+ size 934456
fp16/voices/VARM0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9d4378293b9b32ebfa9f692b1452007efa92f1872d48809326512b71f577a200
3
+ size 705080
fp16/voices/VARM1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e60fe9ec69e744b5f5b1dda2869cf1cde5c34ee83225b4aa53113e19b1a8bdb0
3
+ size 754232
fp16/voices/VARM2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:82ab64e1d9a1c4935f09fab820584d0fc942faea016c8feeae9dcf40099441d0
3
+ size 1131064
fp16/voices/VARM3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c12db9ac957635e8c844895f7dc8866e470db221d8b986bc815bf83b0c01e356
3
+ size 754232
fp16/voices/VARM4.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:36878c0e21496b5ad5122a9b9143e7acce943c931ff0387d52012390e9c6570a
3
+ size 901688