# Diarization Evaluation This document describes the recommended way for partners and users to evaluate Sortformer and Nemotron diarization models using NVIDIA NeMo Speech. Use the NeMo Speech evaluation script as the source of truth: https://github.com/NVIDIA-NeMo/Speech/blob/main/examples/speaker_tasks/diarization/neural_diarizer/e2e_diarize_speech.py The script runs model inference, converts frame-level speaker activity predictions into diarization segments, and reports diarization error rate (DER) using NeMo Speech's diarization scoring utilities. ## Recommended Evaluation Command Install and run from a checkout of NeMo Speech: ```bash git clone https://github.com/NVIDIA-NeMo/Speech.git cd Speech export NEMO_ROOT="$PWD" ``` Evaluate a downloaded `.nemo` checkpoint: ```bash python ${NEMO_ROOT}/examples/speaker_tasks/diarization/neural_diarizer/e2e_diarize_speech.py \ model_path="/path/to/Nemotron-3-Diarization-preview.nemo" \ dataset_manifest="/path/to/diarization_manifest.json" \ batch_size=32 \ collar=0 \ ignore_overlap=false \ precision=bf16 \ compile_encoder=false \ spkcache_len=264 \ chunk_len=340 \ chunk_right_context=40 \ fifo_len=40 \ spkcache_update_period=300 ``` For a Hugging Face model repository, first download the `.nemo` checkpoint and then pass its local path to `model_path`. The input manifest must contain one JSON object per line. For DER evaluation, each record must include an RTTM reference: ```json {"audio_filepath": "/path/to/audio_001.wav", "offset": 0, "duration": 600, "rttm_filepath": "/path/to/audio_001.rttm"} {"audio_filepath": "/path/to/audio_002.wav", "offset": 0, "duration": 580, "rttm_filepath": "/path/to/audio_002.rttm"} ``` If a UEM file is required by the benchmark protocol, include `uem_filepath` in the manifest record. Otherwise, NeMo derives the evaluated region from the reference and hypothesis extents and clamps it to the manifest `offset` and `duration` when those fields are present. ## Metrics ### Diarization Error Rate Diarization Error Rate (DER) is the primary diarization metric. It is the sum of three error components, normalized by the scored reference speaker time: ```text DER = false alarm + missed speech + speaker confusion ``` The NeMo Speech script reports: ```text FA, MISS, CER, DER ``` In the script output, `CER` means speaker confusion error rate, not character error rate. ### Speaker Counting Accuracy Speaker Counting Accuracy (SCA) measures whether the predicted number of speakers exactly matches the reference number of speakers for each recording: ```text SCA = number_of_recordings_with_correct_speaker_count / number_of_recordings ``` Report SCA as a simple percentage: ```text SCA (%) = 100 * SCA ``` NeMo Speech logs this value as `Spk. Count Acc.`. ### Speaker Counting Mean Absolute Error NeMo Speech currently reports Speaker Counting Mean Absolute Error (MAE): ```text MAE = mean(abs(predicted_speaker_count - reference_speaker_count)) ``` The NeMo Speech scorer logs this value as `Spk. Count MAE`. ## Required Reporting Convention When reporting DER for Sortformer or Nemotron diarization models, include the evaluation protocol details with the number. DER is not fully interpretable without these settings. Report at least: | Field | What to report | | --- | --- | | Model | Model name and checkpoint path or revision | | Evaluation code | NeMo Speech repository URL and commit hash | | Dataset | Dataset name, split, and manifest path or release identifier | | Reference annotation | RTTM source and citation, such as paper, dataset URL, or forced-alignment reference | | UEM | Whether UEM regions were used, and the UEM source if applicable | | Overlap | Whether overlapping speech was included in scoring (`ignore_overlap=false`) or excluded (`ignore_overlap=true`) | | Collar | Collar half-width in seconds, for example `0.0` or `0.25` | | Post-processing | Whether post-processing was bypassed or which post-processing YAML was used | | Output resolution | Native or overridden output subsampling factor | | Streaming settings | `spkcache_len`, `chunk_len`, `chunk_left_context`, `chunk_right_context`, `fifo_len`, and `spkcache_update_period` | | Precision | For example `bf16`, `bf16-mixed`, or `32` | | Hardware | GPU model and batch size | Convert NeMo's `DER`, `FA`, `MISS`, `CER`, and `Spk. Count Acc.` rates to percentages by multiplying them by 100; report `Spk. Count MAE` unchanged. Recommended metric table columns: | Dataset | DER (%) | FA (%) | Miss (%) | Confusion (%) | SCA (%) | Speaker Count MAE | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | Example split | 0.00 | 0.00 | 0.00 | 0.00 | 100.00 | 0.0000 | ## Notes on Collar and Overlap The NeMo Speech DER scorer follows NIST `md-eval-22.pl` collar semantics: `collar` is the no-score half-width around each reference boundary. For example, `collar=0.25` excludes 0.25 seconds to the left and 0.25 seconds to the right of each reference boundary. The `ignore_overlap` flag controls whether overlapped speech contributes to DER: ```text ignore_overlap=false # include overlap in DER ignore_overlap=true # exclude overlap from DER ``` For model-card style reporting, prefer the benchmark's official protocol. If the benchmark does not specify a protocol, report DER with overlap included and state the collar value explicitly. ## Where the Metrics Come From in NeMo Speech [e2e_diarize_speech.py](https://github.com/NVIDIA-NeMo/Speech/blob/main/examples/speaker_tasks/diarization/neural_diarizer/e2e_diarize_speech.py) restores the diarization model, sets the test manifest, applies the requested streaming parameters, runs inference, converts predictions to timestamped speaker segments, and calls: ```python score_labels( AUDIO_RTTM_MAP=infer_audio_rttm_dict, all_reference=all_refs, all_hypothesis=all_hyps, all_uem=all_uems, collar=cfg.collar, ignore_overlap=cfg.ignore_overlap, ) ``` The NeMo Speech scoring utility computes cumulative `FA`, `MISS`, speaker confusion (`CER`), and `DER`. It also computes speaker-count accuracy and speaker-count MAE by comparing the number of unique speakers in the reference and hypothesis for each manifest region. ## Stand-alone scoring script [score_diarization.py](https://github.com/NVIDIA-NeMo/Speech/blob/main/scripts/speaker_tasks/score_diarization.py) is a simple user-facing script for diarization scoring. ```bash python ${NEMO_ROOT}/scripts/speaker_tasks/score_diarization.py \ -r $REFERENCE \ -h $HYPOTHESIS \ -c $COLLAR ``` The reference and hypothesis may each be: - A compound RTTM file - A directory containing per-recording RTTM files - A JSON/JSONL diarization manifest When both inputs are directories, RTTM filenames must match exactly. The scorer includes overlapping speech regions (`ignore_overlap=False`).