alexwengg's picture
Upload README.md with huggingface_hub
ed17729 verified
|
Raw
History Blame Contribute Delete
7.08 kB
metadata
license: other
license_name: nvidia-software-and-model-evaluation-license
license_link: >-
  https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-software-and-model-evaluation-license/
base_model: nvidia/Nemotron-3-Diarization-preview
library_name: fluidaudio
pipeline_tag: voice-activity-detection
tags:
  - speaker-diarization
  - streaming
  - coreml
  - apple-silicon
  - ane
  - sortformer

Nemotron 3 Diarization β€” CoreML for Apple Silicon

CoreML conversion of NVIDIA's Nemotron 3 Diarization preview (8-speaker streaming Sortformer, 100M params, 31-layer RoPE Transformer, 10 ms output resolution), optimized for on-device inference on Apple platforms via FluidAudio.

Runs fully on-device β€” Neural Engine or GPU β€” on macOS 14+ / iOS 17+.

Highlights (M5 Pro)

  • Reproduces NVIDIA's model-card accuracy on-device: AMI MHM offline 9.28 DER (card 9.30), verylow 10.13 (card 10.13 β€” exact), AliMeeting Far 9.83 (card 11.14), AliMeeting Near 6.79 (card 7.25) β€” same forced-alignment references, collar 0
  • 8-speaker streaming at 1.04 s latency, ~11 ms ANE per chunk (fast preset)
  • fast128: 9.59 DER at 1004x wall real-time (card protocol, AMI MHM) β€” every FluidAudio preset beats the card's published low-profile number
  • 100% ANE-resident W8A8 split-graph builds at half the weight footprint (c128-split-w8a8: 9.68 DER, 566x wall, ANE only β€” GPU stays free)
  • Beats streaming Sortformer v2 on all 16 AMI test meetings with double the speaker capacity

Which model should I use?

Model Size Audio chunk/call Latency DERΒΉ Per-call cost Pros Cons
low 190 MB 0.72 s 1.04 s 9.90 28 ms ANE / 9.6 ms GPU Best quality at real streaming latency; NVIDIA's reference config (card: 10.35) Heaviest ANE draw per second of audio
fast 190 MB 0.72 s 1.04 s 10.65 10.9 ms ANE / 6.1 ms GPU ~3x cheaper than low at the same latency; leaves ANE room for concurrent ASR Worst DER of the set (+0.75 vs low)
fast32 190 MB 2.56 s 2.88 s 9.94 12.5 ms ANE / 6.2 ms GPU Recommended default β€” low-level quality at near-fast cost, 343x wall 2.88 s latency; not for live-caption UX
fast128 190 MB 10.24 s 10.56 s 9.59 20.6 ms ANE / 7.5 ms GPU Best streaming-lineup DER; 1004x wall 10.5 s latency β€” near-live only
offline 190 MB 27.2 s 30.4 s 9.28 11.2 ms GPU only Highest accuracy (exact card match); ~1680x wall No ANE (compiler limit); 30 s latency
s32_split_w8a8 95 MB 2.56 s 2.88 s 9.86 9.7 ms, 100% ANE Half size, zero GPU use β€” the iOS/battery pick Split mode needs host-side pre-encode (FluidAudio handles it)
c128_split_w8a8 95 MB 10.24 s 10.56 s 9.68 16.1 ms, 100% ANE Batch throughput (566x) without touching the GPU Same split-mode dependency; 10.5 s latency

ΒΉ Card protocol: AMI MHM test, forced-alignment refs, collar 0 (NVIDIA's published numbers under this protocol: 9.30 offline / 10.35 low). Wall RTFx single-stream on M5 Pro. Quick chooser: hard ~1 s latency -> fast (sharing ANE) or low; general use -> fast32; latency-flexible quality -> fast128; Mac batch -> offline; iPhone/iPad or GPU-busy -> the split_w8a8 pair.

Presets

All presets share one checkpoint; they differ in streaming shape. Latency = (chunk + right-context) x 80 ms. DER: card protocol β€” AMI MHM test, 16 meetings, forced-alignment references (nttcslab-sp/diar-forced-alignment), collar 0, overlap included. NVIDIA's card reports 9.30 (offline) / 10.35 (low) under this protocol. Full tables incl. AMI SDM and AliMeeting in BENCHMARKS.md.

Monolithic (single model file, simplest integration)

Preset Latency Audio/call DER (card protocol) ANE/call Use case
low 1.04 s 0.72 s 9.90 28 ms card config (card: 10.35)
fast 1.04 s 0.72 s 10.65 10.9 ms latency-critical + concurrent ASR
fast32 2.88 s 2.56 s 9.94 12.5 ms recommended streaming default
fast128 10.56 s 10.24 s 9.59 20.6 ms near-live throughput, 1004x wall
offline 30.4 s 27.2 s 9.28 GPU only batch (card: 9.30; ~1680x wall)

Other configurations (sub-second latency tiers, intermediate chunk sizes, weight-only int8) were built and benchmarked but are not published β€” they are dominated by the presets above. Open an issue if you need one; regeneration is scripted.

Split-graph (host pre-encode; 100% ANE-resident CoreML subgraph)

Feature stacking, the 1024β†’512 projection (pre_encode_proj_t.bin), state packing, and masks run host-side; the model is the pure-fp transformer+head. FluidAudio handles all of this via splitGraph presets.

Preset Latency DER (card protocol) ANE/call Weights
s32_split_w8a8 2.88 s 9.86 9.7 ms 95 MB β€” streaming on ANE (iOS pick)
c128_split_w8a8 10.56 s 9.68 16.1 ms 95 MB β€” recommended ANE batch

Usage (FluidAudio, Swift)

import FluidAudio

let config = Nemotron3Config.fast32  // or .preset(named: "c128-split-w8a8")!
let models = try await Nemotron3Models.load(
    config: config,
    directory: modelsDirectoryURL  // this repo's files
)
let diarizer = Nemotron3Diarizer(config: config, models: models)

let (probs, frames) = try diarizer.processComplete(audioSamples)  // 16 kHz mono
let segments = Nemotron3Diarizer.segments(probabilities: probs, frameCount: frames)
// -> arrival-ordered speaker segments at 10 ms resolution, up to 8 speakers

Optional VAD gating for sparse audio (2x wall speedup at 55% speech density):

let (probs, frames) = try diarizer.processComplete(audioSamples, speechMask: mask)

Files

  • monolithic/ β€” one .mlmodelc per preset (mel features in, predictions out)
  • split/ β€” W8A8 transformer graphs (packed/attn_bias/output_mask inputs)
  • learnable_sil_emb.bin β€” learned silence embedding, 512 x fp32 (cache compression)
  • pre_encode_proj_t.bin β€” FeatureStacking projection W^T, 1024x512 fp32 (split mode)

All models: fixed shapes, fp16 weights unless noted, minimum_deployment_target iOS 17 / macOS 14. Conversion verified against the NeMo reference at ≀1e-4 single-chunk parity and 99.995% closed-loop frame agreement on real audio.

Attribution & license

Converted from nvidia/Nemotron-3-Diarization-preview and published with NVIDIA's permission; model weights remain subject to NVIDIA's license terms (see license_link β€” currently the evaluation license, to be updated at the model's general release). Conversion pipeline, Swift runtime, and benchmarks by Fluid Inference.

Reference: Streaming Sortformer.