aufklarer's picture
upload fp16 bundle
55a31ab verified
|
Raw History Blame
3.85 kB
metadata
license: other
license_name: nvidia-open-model-license
license_link: https://huggingface.co/nvidia/personaplex-7b-v1/resolve/main/LICENSE
language:
  - en
library_name: onnxruntime
tags:
  - personaplex
  - speech-to-speech
  - full-duplex
  - voice-agent
  - onnx
  - moshi-architecture
base_model: nvidia/personaplex-7b-v1

PersonaPlex 7B ONNX -- fp16 variant

Maximum quality, maximum VRAM. FP16 temporal + FP16 depformer + FP32 mimi.

This is one of four ONNX-quantized bundles of NVIDIA's PersonaPlex 7B -- a full-duplex speech-to-speech model on Kyutai's Moshi architecture. The full collection is at soniqo/PersonaPlex-7B-ONNX.

Quick reference (measured on RTX 5090, VARF2 voice, "helpful" prompt, 50 frames)

This bundle
Disk ~17 GB
Host RAM 1.5 GB
VRAM 18.3 GB
RTF (12.5 Hz) 5.3x (1.0 = realtime)
Quality cos 0.999999 vs FP32 reference

Near-perfect quality. Use when VRAM is plentiful.

All four variants

Variant Disk Host RAM VRAM RTF hidden cos Best for
fp16 ~17 GB 1.5 GB 18.3 GB 5.3x 1.0000 Near-perfect quality ← (this)
mixed ~11 GB 7.9 GB 6.6 GB 3.5x 0.9904 Best quality at low VRAM
int8-nb-dep_gint8 ⭐ ~9.4 GB 1.4 GB 12.1 GB 1.12x 0.9977 Best balance: near realtime RTF, low host RAM, excellent quality
int4-nb-dep_gint8 ~7.6 GB 1.4 GB 9.6 GB 1.12x 0.8774 Smallest viable bundle

The int8-nb-dep_gint8 variant is the recommended ship default: best RTF (1.12x, near realtime), low host RAM (1.4 GB), and excellent quality (cos 0.998).

The mixed variant wins on quality + VRAM combined but its RTF is 3.5x: the INT8 dynamic-quantize pattern adds 384 CPU↔GPU Memcpy bridges which block CUDA Graph capture. Pick this when VRAM is the binding constraint.

How to use

This bundle is consumed by the speech-core C++ inference runtime:

# Download (from speech-core repo root)
PERSONAPLEX_VARIANT=fp16 scripts/download_personaplex_onnx.sh

# Run
build/Release/run_personaplex scripts/personaplex-fp16 50 audio.wav VARF2

Or load directly with ONNX Runtime in any language:

import onnxruntime as ort
sess = ort.InferenceSession("temporal_step.onnx", providers=["CUDAExecutionProvider"])
# inputs: text_token [1,1] int64, audio_tokens [1,16] int64,
#         past_k_all [32,1,32,T,128] float, past_v_all [32,1,32,T,128] float
# outputs: hidden [1,1,4096] float, new_k_all, new_v_all

Files in this variant

File Purpose
mimi_encoder.onnx(+.data) 24 kHz PCM -> 16 audio codebooks @ 12.5 Hz
mimi_decoder.onnx(+.data) 16 audio codebooks @ 12.5 Hz -> 24 kHz PCM
temporal_step.onnx(+.data) One frame of 32-layer 7B temporal transformer, explicit KV-cache I/O
depformer_step.onnx(+.data) One inner step of 6-layer depformer, 16 codebook steps per frame
tokenizer_spm_32k_3.model SentencePiece text tokenizer
voices/<name>.bin 18 voice prompts (NATF*, NATM*, VARF*, VARM*)
system_prompts.bin Pre-tokenized "helpful" / "expert" / "warm" / "direct" prompts
config.json Architecture + precision + measured-metrics metadata

Related

License

Same as the upstream NVIDIA PersonaPlex 7B (NVIDIA Open Model License).