--- license: other license_name: nvidia-open-model-license license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE language: - en library_name: onnxruntime tags: - personaplex - speech-to-speech - full-duplex - voice-agent - onnx - moshi-architecture base_model: nvidia/personaplex-7b-v1 --- # PersonaPlex 7B ONNX -- `int4-nb-dep_gint8` variant Smallest disk. INT4 MatMulNBits temporal (block=32) + custom INT8 depformer + FP32 mimi. Coherent but degraded quality. This is one of four ONNX-quantized bundles of NVIDIA's [PersonaPlex 7B](https://huggingface.co/nvidia/personaplex-7b-v1) -- a full-duplex speech-to-speech model on Kyutai's [Moshi](https://github.com/kyutai-labs/moshi) architecture. The full collection is at [soniqo/PersonaPlex-7B-ONNX](https://huggingface.co/soniqo/PersonaPlex-7B-ONNX). ## Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames) | | This bundle | |---|---| | **Disk** | ~7.6 GB | | **Host RAM** | 1.4 GB | | **VRAM** | 9.6 GB | | **RTF (12.5 Hz)** | 1.12x (1.0 = realtime) | | **Quality** | cos 0.877393 vs FP32 reference | Smallest viable bundle. Visibly degraded but produces coherent English. ## All four variants | Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for | |---|---|---|---|---|---|---| | `fp16` | ~17 GB | 1.5 GB | 18.3 GB | 5.3x | 1.0000 | Near-perfect quality | | `mixed` | ~11 GB | 7.9 GB | 6.6 GB | 3.5x | 0.9904 | Best quality at low VRAM | | `int8-nb-dep_gint8` ⭐ | ~9.4 GB | 1.4 GB | 12.1 GB | 1.12x | 0.9977 | Best balance: near realtime RTF, low host RAM, excellent quality | | `int4-nb-dep_gint8` | ~7.6 GB | 1.4 GB | 9.6 GB | 1.12x | 0.8774 | Smallest viable bundle ← **(this)** | The `int8-nb-dep_gint8` variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998). The `mixed` variant wins on **quality + VRAM combined** but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint. ## How to use This bundle is consumed by the [speech-core](https://github.com/soniqo/speech-core) C++ inference runtime: ```bash # Download (from speech-core repo root) PERSONAPLEX_VARIANT=int4-nb-dep_gint8 scripts/download_personaplex_onnx.sh # Run build/Release/run_personaplex scripts/personaplex-int4-nb-dep_gint8 50 audio.wav VARF2 ``` Or load directly with ONNX Runtime in any language: ```python import onnxruntime as ort sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"]) # inputs: text_token [1,1] int64, audio_tokens [1,16] int64, # past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float # outputs: hidden [1,1,4096] float, new_k_all, new_v_all ``` ## Files in this variant | File | Purpose | |---|---| | `mimi_encoder.onnx`(+`.data`) | 24 kHz PCM -> 16 audio codebooks @ 12.5 Hz | | `mimi_decoder.onnx`(+`.data`) | 16 audio codebooks @ 12.5 Hz -> 24 kHz PCM | | `temporal_step.onnx`(+`.data`) | One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O | | `depformer_step.onnx`(+`.data`) | One inner step of 6-layer depformer, 16 codebook steps per frame | | `tokenizer_spm_32k_3.model` | SentencePiece text tokenizer | | `voices/.bin` | 18 voice prompts (NATF*, NATM*, VARF*, VARM*) | | `system_prompts.bin` | Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts | | `config.json` | Architecture + precision + measured-metrics metadata | ## Related - [soniqo/speech-core](https://github.com/soniqo/speech-core) -- C++ runtime that consumes this bundle, with the `OnnxPersonaPlex` wrapper and CUDA EP routing - [soniqo.audio](https://soniqo.audio) -- the project site ## License Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License).