CAM++ zh+en speaker embedding (ONNX: fp32 / fp16 / int8)
ONNX speaker-embedding model for on-device, CPU-only speaker diarization in a multilingual (Mandarin/Taiwanese + English, code-switched) setting. Repackaged for sherpa-onnx with fp16 and int8 variants added for size/latency comparison.
- Base model: 3D-Speaker
speech_campplus_sv_zh_en_16k-common_advanced(CAM++, D-TDNN), trained on code-switched Mandarin+English. Embedding dim 192, 16 kHz, 80-dim fbank input. - fp32 is the official sherpa-onnx artifact; fp16 / int8 were produced from it with
onnxconverter_common.float16(stats sub-graph kept in fp32) andonnxruntime.quantize_dynamic. - License: Apache-2.0 (inherited from 3D-Speaker / ModelScope).
| file | precision | size |
|---|---|---|
campplus_zh_en_fp32.onnx |
fp32 | 27 MB |
campplus_zh_en_fp16.onnx |
fp16 | 14 MB |
campplus_zh_en_int8.onnx |
int8 (dynamic) | 8.2 MB |
On-device benchmark (Pixel 6, ARM64 CPU, sherpa-onnx, 4 threads)
Per-utterance embedding over real clips; accuracy = speaker separation margin
(mean inter-speaker โ mean intra-speaker cosine distance, higher = better) and clustering outcome;
speed = ms per utterance (5-run average). Baseline = the previous embedder, 3D-Speaker
eres2net_base (zh).
Clean 2-speaker clip (EN + ZH):
| model | size | ms/utt | accuracy (margin) | clustering |
|---|---|---|---|---|
| eres2net_base (old) | 37.8 MB | 256 | 0.715 | 2 spk โ |
| CAM++ fp32 | 27.0 MB | 102 | 0.814 | 2 spk โ |
| CAM++ int8 | 8.2 MB | 287 | 0.738 | 2 spk โ |
| CAM++ fp16 | 14.0 MB | 70 | 0.808 | 2 spk โ |
3-speaker cross-lingual news clip:
| model | size | ms/utt | accuracy (margin) | clustering |
|---|---|---|---|---|
| eres2net_base (old) | 37.8 MB | 316 | 0.275 | 3 spk โ |
| CAM++ fp32 | 27.0 MB | 120 | 0.305 | 3 spk โ |
| CAM++ int8 | 8.2 MB | 330 | 0.284 | 3 spk โ |
| CAM++ fp16 | 14.0 MB | 78 | 0.304 | 3 spk โ |
Takeaways
- fp16 is the best on-device choice: fastest (~1.5ร faster than fp32, ~3.5ร faster than the old eres2net), half the size of fp32, and accuracy indistinguishable from fp32. The Pixel 6's ARMv8.2 fp16 SIMD runs it in hardware.
- int8 is counter-productive on this CPU: slower than fp32 (sherpa-onnx's ORT build has no
optimized
MatMulIntegerkernel for this graph) and slightly less accurate โ it only wins on disk. - CAM++ beats the previous eres2net_base baseline on accuracy, speed, and size at every precision.
Usage (sherpa-onnx)
import sherpa_onnx
ext = sherpa_onnx.SpeakerEmbeddingExtractor(
sherpa_onnx.SpeakerEmbeddingExtractorConfig(
model="campplus_zh_en_fp16.onnx", num_threads=4, provider="cpu"))
Produced for VoxSumDroid (offline Android transcribe + diarize + summarize). Original model ยฉ 3D-Speaker (ModelScope), Apache-2.0.