aufklarer commited on
Commit
96c5eaf
·
verified ·
1 Parent(s): 23e2cd0

upload int4-nb-dep_gint8 bundle

Browse files
.gitattributes CHANGED
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ int4-nb-dep_gint8/depformer_step.onnx.data filter=lfs diff=lfs merge=lfs -text
37
+ int4-nb-dep_gint8/mimi_decoder.onnx.data filter=lfs diff=lfs merge=lfs -text
38
+ int4-nb-dep_gint8/mimi_encoder.onnx.data filter=lfs diff=lfs merge=lfs -text
39
+ int4-nb-dep_gint8/temporal_step.onnx.data filter=lfs diff=lfs merge=lfs -text
int4-nb-dep_gint8/README.md ADDED
@@ -0,0 +1,96 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: nvidia-open-model-license
4
+ license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE
5
+ language:
6
+ - en
7
+ library_name: onnxruntime
8
+ tags:
9
+ - personaplex
10
+ - speech-to-speech
11
+ - full-duplex
12
+ - voice-agent
13
+ - onnx
14
+ - moshi-architecture
15
+ base_model: nvidia/personaplex-7b-v1
16
+ ---
17
+
18
+ # PersonaPlex 7B ONNX -- `int4-nb-dep_gint8` variant
19
+
20
+ Smallest disk. INT4 MatMulNBits temporal (block=32) + custom INT8 depformer + FP32 mimi. Coherent but degraded quality.
21
+
22
+ This is one of four ONNX-quantized bundles of NVIDIA's
23
+ [PersonaPlex 7B](https://huggingface.co/nvidia/personaplex-7b-v1) -- a
24
+ full-duplex speech-to-speech model on Kyutai's
25
+ [Moshi](https://github.com/kyutai-labs/moshi) architecture. The full
26
+ collection is at
27
+ [soniqo/PersonaPlex-7B-ONNX](https://huggingface.co/soniqo/PersonaPlex-7B-ONNX).
28
+
29
+ ## Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames)
30
+
31
+ | | This bundle |
32
+ |---|---|
33
+ | **Disk** | ~7.6 GB |
34
+ | **Host RAM** | 1.4 GB |
35
+ | **VRAM** | 9.6 GB |
36
+ | **RTF (12.5 Hz)** | 1.12x (1.0 = realtime) |
37
+ | **Quality** | cos 0.877393 vs FP32 reference |
38
+
39
+ Smallest viable bundle. Visibly degraded but produces coherent English.
40
+
41
+ ## All four variants
42
+
43
+ | Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for |
44
+ |---|---|---|---|---|---|---|
45
+ | `fp16` | ~17 GB | 1.5 GB | 18.3 GB | 5.3x | 1.0000 | Near-perfect quality |
46
+ | `mixed` | ~11 GB | 7.9 GB | 6.6 GB | 3.5x | 0.9904 | Best quality at low VRAM |
47
+ | `int8-nb-dep_gint8` ⭐ | ~9.4 GB | 1.4 GB | 12.1 GB | 1.12x | 0.9977 | Best balance: near realtime RTF, low host RAM, excellent quality |
48
+ | `int4-nb-dep_gint8` | ~7.6 GB | 1.4 GB | 9.6 GB | 1.12x | 0.8774 | Smallest viable bundle ← **(this)** |
49
+
50
+ The `int8-nb-dep_gint8` variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998).
51
+
52
+ The `mixed` variant wins on **quality + VRAM combined** but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint.
53
+
54
+ ## How to use
55
+
56
+ This bundle is consumed by the [speech-core](https://github.com/soniqo/speech-core) C++ inference runtime:
57
+
58
+ ```bash
59
+ # Download (from speech-core repo root)
60
+ PERSONAPLEX_VARIANT=int4-nb-dep_gint8 scripts/download_personaplex_onnx.sh
61
+
62
+ # Run
63
+ build/Release/run_personaplex scripts/personaplex-int4-nb-dep_gint8 50 audio.wav VARF2
64
+ ```
65
+
66
+ Or load directly with ONNX Runtime in any language:
67
+
68
+ ```python
69
+ import onnxruntime as ort
70
+ sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"])
71
+ # inputs: text_token [1,1] int64, audio_tokens [1,16] int64,
72
+ # past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float
73
+ # outputs: hidden [1,1,4096] float, new_k_all, new_v_all
74
+ ```
75
+
76
+ ## Files in this variant
77
+
78
+ | File | Purpose |
79
+ |---|---|
80
+ | `mimi_encoder.onnx`(+`.data`) | 24 kHz PCM -> 16 audio codebooks @ 12.5 Hz |
81
+ | `mimi_decoder.onnx`(+`.data`) | 16 audio codebooks @ 12.5 Hz -> 24 kHz PCM |
82
+ | `temporal_step.onnx`(+`.data`) | One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O |
83
+ | `depformer_step.onnx`(+`.data`) | One inner step of 6-layer depformer, 16 codebook steps per frame |
84
+ | `tokenizer_spm_32k_3.model` | SentencePiece text tokenizer |
85
+ | `voices/<name>.bin` | 18 voice prompts (NATF*, NATM*, VARF*, VARM*) |
86
+ | `system_prompts.bin` | Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts |
87
+ | `config.json` | Architecture + precision + measured-metrics metadata |
88
+
89
+ ## Related
90
+
91
+ - [soniqo/speech-core](https://github.com/soniqo/speech-core) -- C++ runtime that consumes this bundle, with the `OnnxPersonaPlex` wrapper and CUDA EP routing
92
+ - [soniqo.audio](https://soniqo.audio) -- the project site
93
+
94
+ ## License
95
+
96
+ Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License).
int4-nb-dep_gint8/config.json ADDED
@@ -0,0 +1,125 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "personaplex",
3
+ "version": "personaplex-7b-v1",
4
+ "base_model": "nvidia/personaplex-7b-v1",
5
+ "bundle_variant": "int4-nb-dep_gint8",
6
+ "bundle_size_gb": 7.6,
7
+ "precision": {
8
+ "temporal": "int4_matmulnbits_b32",
9
+ "depformer": "custom_int8_gather",
10
+ "mimi": "fp32"
11
+ },
12
+ "metrics": {
13
+ "host_ram_gb": 1.4,
14
+ "vram_gb": 9.6,
15
+ "rtf": 1.12,
16
+ "hidden_cos_vs_fp32": 0.877393,
17
+ "verdict": "Smallest viable bundle. Visibly degraded but produces coherent English."
18
+ },
19
+ "description": "Smallest disk. INT4 MatMulNBits temporal (block=32) + custom INT8 depformer + FP32 mimi. Coherent but degraded quality.",
20
+ "temporal": {
21
+ "dim": 4096,
22
+ "num_layers": 32,
23
+ "num_heads": 32,
24
+ "head_dim": 128,
25
+ "hidden_scale": 4.125,
26
+ "n_q": 16,
27
+ "card": 2048,
28
+ "text_card": 32000,
29
+ "context": 3000,
30
+ "max_period": 10000
31
+ },
32
+ "depformer": {
33
+ "dim": 1024,
34
+ "num_layers": 6,
35
+ "num_heads": 16,
36
+ "head_dim": 64,
37
+ "dim_feedforward": 4224,
38
+ "num_steps": 16,
39
+ "card": 2048,
40
+ "context": 8,
41
+ "weights_per_step": true,
42
+ "multi_linear": true,
43
+ "pos_emb": "none"
44
+ },
45
+ "mimi": {
46
+ "sample_rate": 24000,
47
+ "frame_rate": 12.5,
48
+ "num_codebooks_used": 8,
49
+ "num_codebooks_internal": 32,
50
+ "codebook_size": 2048,
51
+ "samples_per_frame": 1920
52
+ },
53
+ "delays": [
54
+ 0,
55
+ 0,
56
+ 1,
57
+ 1,
58
+ 1,
59
+ 1,
60
+ 1,
61
+ 1,
62
+ 1,
63
+ 0,
64
+ 1,
65
+ 1,
66
+ 1,
67
+ 1,
68
+ 1,
69
+ 1,
70
+ 1
71
+ ],
72
+ "sampling": {
73
+ "audio_temp": 0.8,
74
+ "audio_top_k": 250,
75
+ "text_temp": 0.7,
76
+ "text_top_k": 25
77
+ },
78
+ "voices": [
79
+ "NATF0",
80
+ "NATF1",
81
+ "NATF2",
82
+ "NATF3",
83
+ "NATM0",
84
+ "NATM1",
85
+ "NATM2",
86
+ "NATM3",
87
+ "VARF0",
88
+ "VARF1",
89
+ "VARF2",
90
+ "VARF3",
91
+ "VARF4",
92
+ "VARM0",
93
+ "VARM1",
94
+ "VARM2",
95
+ "VARM3",
96
+ "VARM4"
97
+ ],
98
+ "system_prompts": [
99
+ "helpful",
100
+ "expert",
101
+ "warm",
102
+ "direct"
103
+ ],
104
+ "voice_embedding": {
105
+ "frames": 50,
106
+ "dim": 4096,
107
+ "default_scale": 10.0
108
+ },
109
+ "prefill_layout": {
110
+ "voice_embedding_prefix_frames": 50,
111
+ "voice_cache_replay_frames": 4,
112
+ "silence_spacer_frames": 6
113
+ },
114
+ "graphs": {
115
+ "mimi_encoder": "mimi_encoder.onnx",
116
+ "mimi_decoder": "mimi_decoder.onnx",
117
+ "temporal_step": "temporal_step.onnx",
118
+ "depformer_step": "depformer_step.onnx"
119
+ },
120
+ "auxiliary": {
121
+ "tokenizer": "tokenizer_spm_32k_3.model",
122
+ "system_prompts": "system_prompts.bin",
123
+ "voices_dir": "voices/"
124
+ }
125
+ }
int4-nb-dep_gint8/depformer_step.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0a29b20d20fdbc80d52b0f21ab58af51a904a0b8ae00d8d1727404c2d0b0ba9d
3
+ size 73941
int4-nb-dep_gint8/depformer_step.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:45ec0c1930f410c4252c22ba7d1db2ecf659b145c26d6deb269643c99d59f4be
3
+ size 1499036672
int4-nb-dep_gint8/mimi_decoder.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0fd6447a47893f7460a852a1178fe6a537301709bfc8d7cae5dc033bfb2bef96
3
+ size 67748881
int4-nb-dep_gint8/mimi_decoder.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9822aa363c73101d58eaaa7ce1d33e5e953576be8823cd0c2fecf2dbacef0bb8
3
+ size 160713728
int4-nb-dep_gint8/mimi_encoder.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f70c97e6aa4194403f4de76b3eaafd47913b942e1de4154f34b82aa9d34bb901
3
+ size 763072
int4-nb-dep_gint8/mimi_encoder.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:16d14e7b15815b546de821d0adf75082254095e9db0c1e3278745aa1bbfdecf9
3
+ size 223886080
int4-nb-dep_gint8/system_prompts.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:00ddeee95bd65c5dd6b9b75004546fd5ac595885acff49fb72ed0dd1134ce6ad
3
+ size 223
int4-nb-dep_gint8/temporal_step.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c3964f16adad7d4a63d3479e51f69c4bb278c771974c74cf793991beeefcf5d1
3
+ size 524766
int4-nb-dep_gint8/temporal_step.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:17d6c4452b45ec8dac69684df98ab4d7a8ca4d6d3e16e6c54597dae288f345d2
3
+ size 5172920320
int4-nb-dep_gint8/tokenizer_spm_32k_3.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:78d4336533ddc26f9acf7250d7fb83492152196c6ea4212c841df76933f18d2d
3
+ size 552778
int4-nb-dep_gint8/voices/NATF0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fd7e4c1e37b3dc5a5583c9ee0cd66a9c1e64a86cdfabcd0bed1d68bf93e8d61c
3
+ size 836152
int4-nb-dep_gint8/voices/NATF1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9e885da22e8eb20b0cd2be99eb08cbc73fba78c73a7e0ecddb6359d94c7dea95
3
+ size 787000
int4-nb-dep_gint8/voices/NATF2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b38778578029eaebb6d0a69b0b5dbb150b058fb7f7c7f9070fcf967d22a8ecbc
3
+ size 836152
int4-nb-dep_gint8/voices/NATF3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:247d0ef94f463090bd7e2ba83d1aa475f2baf55b7d8cd09e95eb02eef22abcd5
3
+ size 836152
int4-nb-dep_gint8/voices/NATM0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5a73ed98b4d7e0012b496d9859b16ba6a5eb9672f4f7bfad0253d25df5a61510
3
+ size 819768
int4-nb-dep_gint8/voices/NATM1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c52210a631f20b378537e60ed1eeff3b9930b62faf0083249e4979d7c516102b
3
+ size 836152
int4-nb-dep_gint8/voices/NATM2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ee7082e5fd5a11aab45109c7c20c4042bd87237cec413b83893bba2a1d765434
3
+ size 819768
int4-nb-dep_gint8/voices/NATM3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:84b93a791b440862a9bea2c5c6a17b34fb2d595a10173a29195207112676258c
3
+ size 754232
int4-nb-dep_gint8/voices/VARF0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6ae429e01a3e0d03313dc1749e6182a6360cb58909627673a3569e2669015f03
3
+ size 1114680
int4-nb-dep_gint8/voices/VARF1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cebc82dcd2040250e767ba3d037a9e6b7cee5a5c07b66c939eb866197ff9aac0
3
+ size 819768
int4-nb-dep_gint8/voices/VARF2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e10657ee1dde618a419c1d2ef10cb4677b02b7d5335cd90b82265a2ee37475d5
3
+ size 836152
int4-nb-dep_gint8/voices/VARF3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:14edf219fdd46edf01a71b5f6c7ed06dfc0563bbaa920cb38002298d18aec816
3
+ size 1016376
int4-nb-dep_gint8/voices/VARF4.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e8a095738ef56ebd3ea29c990c3f8a901b5748874d6538d3410a1a62e6a9e0ce
3
+ size 934456
int4-nb-dep_gint8/voices/VARM0.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9d4378293b9b32ebfa9f692b1452007efa92f1872d48809326512b71f577a200
3
+ size 705080
int4-nb-dep_gint8/voices/VARM1.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e60fe9ec69e744b5f5b1dda2869cf1cde5c34ee83225b4aa53113e19b1a8bdb0
3
+ size 754232
int4-nb-dep_gint8/voices/VARM2.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:82ab64e1d9a1c4935f09fab820584d0fc942faea016c8feeae9dcf40099441d0
3
+ size 1131064
int4-nb-dep_gint8/voices/VARM3.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c12db9ac957635e8c844895f7dc8866e470db221d8b986bc815bf83b0c01e356
3
+ size 754232
int4-nb-dep_gint8/voices/VARM4.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:36878c0e21496b5ad5122a9b9143e7acce943c931ff0387d52012390e9c6570a
3
+ size 901688