CAM++ zh+en speaker embedding (ONNX: fp32 / fp16 / int8)

ONNX speaker-embedding model for on-device, CPU-only speaker diarization in a multilingual (Mandarin/Taiwanese + English, code-switched) setting. Repackaged for sherpa-onnx with fp16 and int8 variants added for size/latency comparison.

  • Base model: 3D-Speaker speech_campplus_sv_zh_en_16k-common_advanced (CAM++, D-TDNN), trained on code-switched Mandarin+English. Embedding dim 192, 16 kHz, 80-dim fbank input.
  • fp32 is the official sherpa-onnx artifact; fp16 / int8 were produced from it with onnxconverter_common.float16 (stats sub-graph kept in fp32) and onnxruntime.quantize_dynamic.
  • License: Apache-2.0 (inherited from 3D-Speaker / ModelScope).
file precision size
campplus_zh_en_fp32.onnx fp32 27 MB
campplus_zh_en_fp16.onnx fp16 14 MB
campplus_zh_en_int8.onnx int8 (dynamic) 8.2 MB

On-device benchmark (Pixel 6, ARM64 CPU, sherpa-onnx, 4 threads)

Per-utterance embedding over real clips; accuracy = speaker separation margin (mean inter-speaker โˆ’ mean intra-speaker cosine distance, higher = better) and clustering outcome; speed = ms per utterance (5-run average). Baseline = the previous embedder, 3D-Speaker eres2net_base (zh).

Clean 2-speaker clip (EN + ZH):

model size ms/utt accuracy (margin) clustering
eres2net_base (old) 37.8 MB 256 0.715 2 spk โœ“
CAM++ fp32 27.0 MB 102 0.814 2 spk โœ“
CAM++ int8 8.2 MB 287 0.738 2 spk โœ“
CAM++ fp16 14.0 MB 70 0.808 2 spk โœ“

3-speaker cross-lingual news clip:

model size ms/utt accuracy (margin) clustering
eres2net_base (old) 37.8 MB 316 0.275 3 spk โœ“
CAM++ fp32 27.0 MB 120 0.305 3 spk โœ“
CAM++ int8 8.2 MB 330 0.284 3 spk โœ“
CAM++ fp16 14.0 MB 78 0.304 3 spk โœ“

Takeaways

  • fp16 is the best on-device choice: fastest (~1.5ร— faster than fp32, ~3.5ร— faster than the old eres2net), half the size of fp32, and accuracy indistinguishable from fp32. The Pixel 6's ARMv8.2 fp16 SIMD runs it in hardware.
  • int8 is counter-productive on this CPU: slower than fp32 (sherpa-onnx's ORT build has no optimized MatMulInteger kernel for this graph) and slightly less accurate โ€” it only wins on disk.
  • CAM++ beats the previous eres2net_base baseline on accuracy, speed, and size at every precision.

Usage (sherpa-onnx)

import sherpa_onnx
ext = sherpa_onnx.SpeakerEmbeddingExtractor(
    sherpa_onnx.SpeakerEmbeddingExtractorConfig(
        model="campplus_zh_en_fp16.onnx", num_threads=4, provider="cpu"))

Produced for VoxSumDroid (offline Android transcribe + diarize + summarize). Original model ยฉ 3D-Speaker (ModelScope), Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support