|
Download int4-nb-dep_gint8/README.md from soniqo/PersonaPlex-7B-ONNX: direct link, hf CLI and curl.
- Browser
- Download file 3.96 kB
-
https://huggingface.co/soniqo/PersonaPlex-7B-ONNX/resolve/55a31aba70f456e532a988b3af0db1148b3e0220/int4-nb-dep_gint8/README.md
- Command line
-
hf download hf://soniqo/PersonaPlex-7B-ONNX@55a31aba70f456e532a988b3af0db1148b3e0220/int4-nb-dep_gint8/README.md
-
curl -L -o README.md https://huggingface.co/soniqo/PersonaPlex-7B-ONNX/resolve/55a31aba70f456e532a988b3af0db1148b3e0220/int4-nb-dep_gint8/README.md
3.96 kB
| license: other | |
| license_name: nvidia-open-model-license | |
| license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE | |
| language: | |
| - en | |
| library_name: onnxruntime | |
| tags: | |
| - personaplex | |
| - speech-to-speech | |
| - full-duplex | |
| - voice-agent | |
| - onnx | |
| - moshi-architecture | |
| base_model: nvidia/personaplex-7b-v1 | |
| # PersonaPlex 7B ONNX -- `int4-nb-dep_gint8` variant | |
| Smallest disk. INT4 MatMulNBits temporal (block=32) + custom INT8 depformer + FP32 mimi. Coherent but degraded quality. | |
| This is one of four ONNX-quantized bundles of NVIDIA's | |
| [PersonaPlex 7B](https://huggingface.co/nvidia/personaplex-7b-v1) -- a | |
| full-duplex speech-to-speech model on Kyutai's | |
| [Moshi](https://github.com/kyutai-labs/moshi) architecture. The full | |
| collection is at | |
| [soniqo/PersonaPlex-7B-ONNX](https://huggingface.co/soniqo/PersonaPlex-7B-ONNX). | |
| ## Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames) | |
| | | This bundle | | |
| |---|---| | |
| | **Disk** | ~7.6 GB | | |
| | **Host RAM** | 1.4 GB | | |
| | **VRAM** | 9.6 GB | | |
| | **RTF (12.5 Hz)** | 1.12x (1.0 = realtime) | | |
| | **Quality** | cos 0.877393 vs FP32 reference | | |
| Smallest viable bundle. Visibly degraded but produces coherent English. | |
| ## All four variants | |
| | Variant | Disk | Host RAM | VRAM | RTF | hidden cos | Best for | | |
| |---|---|---|---|---|---|---| | |
| | `fp16` | ~17 GB | 1.5 GB | 18.3 GB | 5.3x | 1.0000 | Near-perfect quality | | |
| | `mixed` | ~11 GB | 7.9 GB | 6.6 GB | 3.5x | 0.9904 | Best quality at low VRAM | | |
| | `int8-nb-dep_gint8` ⭐ | ~9.4 GB | 1.4 GB | 12.1 GB | 1.12x | 0.9977 | Best balance: near realtime RTF, low host RAM, excellent quality | | |
| | `int4-nb-dep_gint8` | ~7.6 GB | 1.4 GB | 9.6 GB | 1.12x | 0.8774 | Smallest viable bundle ← **(this)** | | |
| The `int8-nb-dep_gint8` variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998). | |
| The `mixed` variant wins on **quality + VRAM combined** but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint. | |
| ## How to use | |
| This bundle is consumed by the [speech-core](https://github.com/soniqo/speech-core) C++ inference runtime: | |
| ```bash | |
| # Download (from speech-core repo root) | |
| PERSONAPLEX_VARIANT=int4-nb-dep_gint8 scripts/download_personaplex_onnx.sh | |
| # Run | |
| build/Release/run_personaplex scripts/personaplex-int4-nb-dep_gint8 50 audio.wav VARF2 | |
| ``` | |
| Or load directly with ONNX Runtime in any language: | |
| ```python | |
| import onnxruntime as ort | |
| sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"]) | |
| # inputs: text_token [1,1] int64, audio_tokens [1,16] int64, | |
| # past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float | |
| # outputs: hidden [1,1,4096] float, new_k_all, new_v_all | |
| ``` | |
| ## Files in this variant | |
| | File | Purpose | | |
| |---|---| | |
| | `mimi_encoder.onnx`(+`.data`) | 24 kHz PCM -> 16 audio codebooks @ 12.5 Hz | | |
| | `mimi_decoder.onnx`(+`.data`) | 16 audio codebooks @ 12.5 Hz -> 24 kHz PCM | | |
| | `temporal_step.onnx`(+`.data`) | One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O | | |
| | `depformer_step.onnx`(+`.data`) | One inner step of 6-layer depformer, 16 codebook steps per frame | | |
| | `tokenizer_spm_32k_3.model` | SentencePiece text tokenizer | | |
| | `voices/<name>.bin` | 18 voice prompts (NATF*, NATM*, VARF*, VARM*) | | |
| | `system_prompts.bin` | Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts | | |
| | `config.json` | Architecture + precision + measured-metrics metadata | | |
| ## Related | |
| - [soniqo/speech-core](https://github.com/soniqo/speech-core) -- C++ runtime that consumes this bundle, with the `OnnxPersonaPlex` wrapper and CUDA EP routing | |
| - [soniqo.audio](https://soniqo.audio) -- the project site | |
| ## License | |
| Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License). | |