Audio Classification
Transformers
ONNX
Safetensors
English
dualturn_endpointing
feature-extraction
turn-taking
endpointing
end-of-turn
voice-activity-detection
voice-agents
conversation
speech
audio
mimi
dualturn
real-time
custom_code
Instructions to use anyreach-ai/dualturn-endpointing with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anyreach-ai/dualturn-endpointing with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="anyreach-ai/dualturn-endpointing", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anyreach-ai/dualturn-endpointing", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 16,926 Bytes
d78eb02 36bb752 d78eb02 36bb752 d78eb02 36bb752 0382386 36bb752 d78eb02 36bb752 964b328 36bb752 964b328 36bb752 964b328 0382386 36bb752 7bf221f 36bb752 964b328 0382386 213dae6 0382386 c3860ed 0382386 bf7bc8e c3860ed 0382386 bf7bc8e 0382386 bf7bc8e 0382386 213dae6 bf7bc8e 213dae6 bf7bc8e 213dae6 bf7bc8e 0382386 7bf221f 36bb752 7bf221f 36bb752 7bf221f 964b328 36bb752 7bf221f 964b328 36bb752 964b328 36bb752 964b328 36bb752 0382386 964b328 0382386 964b328 36bb752 964b328 36bb752 964b328 36bb752 964b328 36bb752 964b328 36bb752 964b328 36bb752 0382386 213dae6 0382386 bf7bc8e c3860ed bf7bc8e 36bb752 0382386 36bb752 964b328 0382386 36bb752 0382386 36bb752 0382386 c3860ed bf7bc8e 36bb752 0382386 36bb752 7bf221f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 | ---
license: apache-2.0
library_name: transformers
pipeline_tag: audio-classification
language:
- en
base_model:
- kyutai/mimi
tags:
- turn-taking
- endpointing
- end-of-turn
- voice-activity-detection
- voice-agents
- conversation
- speech
- audio
- mimi
- dualturn
- onnx
- real-time
arxiv: 2603.08216
---
# DualTurn Endpointing
A small, **causal, audio-only** turn-taking model for voice agents. It listens to **both channels** of a
live conversation (user + agent) and emits, frame-by-frame at 12.5 Hz, the signals a voice agent needs to
decide when to speak β with **end-of-turn (endpointing)** as the primary output.
Adapted from the **DualTurn** paper ([arXiv:2603.08216](https://arxiv.org/abs/2603.08216)) and fine-tuned
for the **endpointing head**. It runs on frozen [Mimi](https://huggingface.co/kyutai/mimi) features, adds
only **~1.5M parameters** on top of the encoder, needs **no ASR and no cloud**, and is designed to run
**on-device / on CPU**.
## How it's meant to be used
The model **sits inside a live two-channel call** between a **user** (CH0) and the **agent** (CH1). It
continuously listens to **both** channels at the same time and, every 80 ms, emits per-frame predictions:
**VAD** and **FVAD** for *both* speakers, and the **user's end-of-turn (EOT)**. A voice agent consumes
these β chiefly the user's EOT β to decide *when it's the agent's turn to speak* (and VAD/FVAD to sense
who's talking and who's about to talk). The agent channel is not just context: hearing the agent's own
speech (and overlaps/interruptions) is what makes the user-EOT judgment reliable in a real conversation.
It can also run as a **single-stream (user-only) detector**: feed the user audio and **silence the agent
channel** (pass mono, or a zeros agent channel). You still get the user's EOT + user VAD/FVAD; you simply
lose the agent-context benefit. This is the setup used for benchmarking and for user-mic-only deployments.
| Output | Shape | Channels | Meaning |
|---|---|---|---|
| `eot_probs` | `[1, T, 1]` | **user** | P(user has finished their turn) β the endpointing signal |
| `vad_probs` | `[1, T, 2]` | user, agent | P(speaking) |
| `fvad_probs` | `[1, T, 2, 4]` | user, agent | P(channel speaks within a future window) β 4 horizons: 0β240 / 240β640 / 640β1200 / 1200β2000 ms |
Frames are **80 ms (12.5 Hz)**; `T β audio_seconds Γ 12.5`. Streaming/causal β outputs at time *t* use
only audio up to *t*.
## Inference modes & latency
Four ready-to-use paths (all **fp32, same weights**). **Offline** processes a whole file at once;
**streaming** runs per 80 ms tick for live calls. ONNX paths need only `onnxruntime` (no PyTorch).
| # | mode | runtime | latency (CPU, 4 threads) | deps |
|---|---|---|---|---|
| 1 | Offline | PyTorch β `model(wav, sr)` | ~7.5 ms/frame (batch, ~11Γ real-time) | torch + transformers |
| 2 | Streaming | PyTorch β `model.streaming()` | **~63 ms / 80 ms tick** (real-time, flat) | torch + transformers |
| 3 | Offline | ONNX β `model.onnx` | ~7.3 ms/frame (~11Γ real-time) | onnxruntime |
| 4 | Streaming | ONNX β `stream_tick.onnx` | **~52 ms / 80 ms tick** (real-time, flat, tightest tail) | onnxruntime |
*Offline latency is amortized batch throughput (all frames returned after the file is processed); streaming
latency is the real per-tick cost in a live call β both well under the 80 ms budget. Streaming is causal +
bounded (transformer KV capped at 250 frames = 10 s), so per-tick latency stays flat over arbitrarily long
calls. Streaming outputs are decision-equivalent to offline β they match the batch model to within ~0.02β0.03
(fp32 op-order drift in the incremental path), which doesn't move threshold+sustain endpoint decisions.*
### Inference interval (compute β latency)
Two streaming ONNX graphs are shipped β pick by how often you need to react vs how much CPU you want to
spend. **Both are strictly causal** (the transformer runs frame-by-frame *inside* the graph β no future
audio) and **decision-equivalent to offline** β they track the batch model to within ~0.02β0.03 (fp32
op-order drift in the streaming path; the two graphs are bit-identical to *each other*, Ξβ0). The 240 ms
graph runs one shared encoder pass over 3 frames instead of three separate passes, so it's ~1.8Γ cheaper
per second of audio:
- **`stream_tick.onnx`** β 1 frame per call (`[2, 1920]` = 80 ms).
- **`stream_tick_240.onnx`** β 3 frames per call (`[2, 5760]` = 240 ms), returns all 3 Γ 80 ms frames.
Measured on a real call (CPU, 4 threads, carried KV cache). **Grace** = idle time left in the interval after
inference (`interval β latency`); **CPU busy** = per-stream duty cycle, fraction of wall-clock the inference
runs to keep up with one live stream (`latency Γ calls/sec`):
| graph | infer every | latency / call | calls / audio-sec | grace / call | CPU busy / stream |
|---|---|---|---|---|---|
| `stream_tick.onnx` | **80 ms** | 52 ms | 12.5 | 28 ms | ~65% |
| `stream_tick_240.onnx` | **240 ms** | 85 ms | 4.17 | 155 ms | ~36% |
- **80 ms β snappiest, but heaviest.** Each call eats 52 ms β only **28 ms grace**, pool busy **~65%** per
stream. Fastest reactions, least room for anything else.
- **240 ms β much cheaper, lots of slack.** Each call takes 85 ms (vs 3 Γ 52 = 156 ms for three 80 ms ticks)
but leaves **155 ms grace** and sits at **~36% busy**.
- **The real win is at scale.** CPU per second of audio drops **~1.8Γ** (650 β 354 ms/s, β45% less), so **one
machine holds roughly ~1.8Γ as many concurrent calls** β the per-stream % looks modest, but that factor
multiplies across every simultaneous call in your fleet. *(`CPU busy %` is a single-stream duty cycle
measured with 4 threads β read it as relative cost between graphs, not absolute single-core load.)*
- **Accuracy is unaffected** β the 240 ms graph is decision-equivalent to the 80 ms one (endpoint decisions
match); the only cost is that you learn the turn ended at the next 240 ms boundary (up to one interval later).
## 1. Offline β PyTorch
```python
import torch, torchaudio
from transformers import AutoModel
model = AutoModel.from_pretrained("anyreach-ai/dualturn-endpointing", trust_remote_code=True).eval()
wav, sr = torchaudio.load("conversation.wav") # [2, T] CH0 = user, CH1 = agent
with torch.no_grad():
out = model(wav, sr=sr)
out.eot_probs # [1, T, 1] P(user end-of-turn)
out.vad_probs # [1, T, 2] P(speaking) β (user, agent)
out.fvad_probs # [1, T, 2, 4] P(speaks soon) β (user, agent) Γ (0-240/240-640/640-1200/1200-2000 ms)
```
Turn the end-of-turn stream into a decision with a thin policy β e.g. fire when `eot_probs` stays above a
threshold for a short sustain window:
```python
eot = out.eot_probs[0, :, 0]
thr, sustain = 0.5, 3 # 3 frames β 240 ms
fired = (eot > thr).unfold(0, sustain, 1).all(-1).nonzero() # frame indices of detected turn-ends
```
Mono input `[T]` or `[1, T]` is also accepted (treated as user with a silent agent β a single-stream
fallback). Any sample rate works (`sr=`); audio is internally resampled to **24 kHz** (Mimi's input rate).
The call above is the **offline** mode (whole file at once). For live calls, use the **streaming** mode below.
## 2. Streaming β PyTorch (real-time)
For live calls, don't re-run the whole file each tick β use the stateful streamer. It's **causal** (uses
only past audio), keeps **bounded state** (transformer KV window + LSTM state), and runs **~63 ms per
80 ms tick on CPU** (β₯4 threads), **flat regardless of call length**.
```python
from transformers import AutoModel
model = AutoModel.from_pretrained("anyreach-ai/dualturn-endpointing", trust_remote_code=True).eval()
streamer = model.streaming()
# feed audio as it arrives, in 80 ms chunks: [2, 1920] @ 24 kHz (CH0=user, CH1=agent)
for chunk in chunks_at_24kHz: # mono [1920] is also accepted (silent agent)
out = streamer.push(chunk) # None until the next 12.5 Hz frame is ready
if out is not None:
p_eot = out.eot_probs[0, -1, 0].item() # latest P(user end-of-turn)
# fire your end-of-turn policy on p_eot (threshold + short sustain)
streamer.reset() # between calls
```
`push` returns a `DualTurnOutput` (same `eot_probs` / `vad_probs` / `fvad_probs`) for the new frame(s).
It matches the offline model to ~3e-4 mean H1. Feed **24 kHz** (resample the continuous stream once, not
per-chunk, to avoid chunk-edge artifacts).
## 3. Offline β ONNX (no PyTorch)
A self-contained `model.onnx` (Mimi encoder + readout in one graph) is included for CPU / on-device
deployment with just `onnxruntime`. Input is a **2-channel waveform at 24 kHz** `[2, T]` (CH0=user,
CH1=agent); resample first with any tool.
```python
import numpy as np, librosa, onnxruntime as ort
from huggingface_hub import hf_hub_download
onnx_path = hf_hub_download("anyreach-ai/dualturn-endpointing", "model.onnx")
sess = ort.InferenceSession(onnx_path, providers=["CPUExecutionProvider"])
wav, sr = librosa.load("conversation.wav", sr=24000, mono=False) # [2, T] @ 24 kHz, CH0=user CH1=agent
eot, vad, fvad = sess.run(None, {"audio_24k": wav.astype("float32")})
# eot [1, T', 1] P(user end-of-turn)
# vad [1, T', 2] P(speaking) β (user, agent)
# fvad [1, T', 2, 4] P(speaks soon) β (user, agent) Γ (0-240/240-640/640-1200/1200-2000 ms)
eot = eot[0, :, 0]
fired = np.where(np.convolve((eot > 0.5).astype(int), np.ones(3), "valid") == 3)[0] # sustained β₯240 ms
```
Runs comfortably faster than real-time on CPU (β1 s of compute per 10 s of audio, single-threaded). The
graph is fully self-contained β no `torch` or `transformers` needed at inference.
## 4. Streaming β ONNX (real-time, no PyTorch)
Torch-free real-time streaming via `stream_tick.onnx` + the `onnx_streaming.py` helper (`DualTurnONNXStreamer`).
Each 80 ms tick only computes the **new frame** β the graph carries all state (conv buffer, transformer **KV
window**, downsample history, LSTM state) in/out, so latency is low (**~52 ms/tick**) and **flat over long
calls**. Decision-equivalent to the PyTorch model.
```python
from onnx_streaming import DualTurnONNXStreamer # download onnx_streaming.py + stream_tick.onnx from the repo
streamer = DualTurnONNXStreamer() # auto-downloads stream_tick.onnx (via huggingface_hub)
for chunk in chunks_24k: # each: [2, 1920] float32 @ 24 kHz (CH0=user, CH1=agent; mono OK)
out = streamer.push(chunk)
p_eot = out["eot"] # P(user end-of-turn); also out["vad"] [2], out["fvad"] [2,4]
# fire your end-of-turn policy on p_eot (threshold + short sustain)
streamer.reset() # between calls
```
Needs only `onnxruntime` + `numpy` (+ `huggingface_hub` for the auto-download). Feed **24 kHz** audio in
**exactly 1920-sample (80 ms)** chunks.
To spend ~1.8Γ less CPU (at the cost of reacting every 240 ms instead of 80 ms), use the **240 ms graph**
`stream_tick_240.onnx` via `DualTurnONNXStreamer240` β same helper file, `[2, 5760]` chunks, returns the 3
Γ 80 ms frames per call. Both graphs are strictly causal and decision-equivalent to offline; see [Inference interval](#inference-interval-compute--latency).
```python
from onnx_streaming import DualTurnONNXStreamer240 # downloads stream_tick_240.onnx
streamer = DualTurnONNXStreamer240()
for chunk in chunks_240ms: # each: [2, 5760] float32 @ 24 kHz
for fr in streamer.push(chunk): # list of 3 frames: fr["eot"], fr["vad"], fr["fvad"]
... # apply your end-of-turn policy on fr["eot"]
streamer.reset()
```
## Performance
Benchmarked on **LiveKit's [eot-bench](https://huggingface.co/datasets/livekit/eot-bench-data)** (a neutral
end-of-turn benchmark). Metric is a latency / false-cutoff Pareto: **FC@N ms** = % of mid-turn pauses
falsely fired on at an N-ms latency budget (**lower is better**); **Lat@X%** = ms of dead-air after a true
turn-end at an X% false-cutoff budget (**lower is better**). Our harness reproduces LiveKit's published
SmartTurn numbers exactly, so rows are directly comparable; DualTurn is scored single-stream
(agent = silence) β an honest handicap. English slice:
| # | Model | Params | FC@300 β | FC@600 β | Lat@5% β | Lat@10% β | Runs on |
|---|---|---|---:|---:|---:|---:|---|
| 1 | LiveKit Turn Detector v1 | undisclosed | 9.9% | 4.5% | 543 ms | 295 ms | βοΈ Cloud |
| 2 | Deepgram Flux | undisclosed | 12.9% | 9.9% | 1151 ms | 548 ms | βοΈ Cloud |
| **3** | **DualTurn Endpointing (this model)** | **~1.5M** | **21.8%** | **10.2%** | 1115 ms | 610 ms | β
**CPU / GPU** |
| 4 | ultraVAD (Ultravox / Llama-8B) | ~8B | 27.7% | 11.9% | 899 ms | 663 ms | π₯οΈ GPU |
| 5 | LiveKit Turn Detector v1-mini | undisclosed | 27.8% | 12.1% | 1070 ms | 698 ms | β
CPU |
| 6 | SmartTurn v3.2 | ~8M | 35.2% | 14.8% | 1051 ms | 739 ms | β
CPU |
| 7 | AssemblyAI | undisclosed | 49.4% | 14.6% | 1049 ms | 713 ms | βοΈ Cloud |
| 8 | VAD baseline (Silero) | ~0.3β1.5M | 55.6% | 21.7% | 1600 ms | 1000 ms | β
CPU |
**DualTurn Endpointing is the best on-device / audio-only turn detector on the board** β beating LiveKit's
own on-device v1-mini, SmartTurn v3.2, and the open 8B ultraVAD, at ~1.5M params without transcription or
cloud. Only closed cloud end-of-turn services lead. It holds the same #1-on-device rank on **zero-shot
Spanish** (not trained on Spanish), evidence that turn-taking here is language-agnostic acoustics.
## How it works
Adapted from the DualTurn paper ([arXiv:2603.08216](https://arxiv.org/abs/2603.08216)) and fine-tuned for endpointing: both channels are encoded by a **frozen
Mimi** speech encoder into 12.5 Hz continuous features; a small **causal** readout (~1.5M params) produces
the per-frame turn-taking probabilities above. Audio-only (no ASR), streaming, and light enough for CPU /
on-device inference. The Mimi encoder is downloaded automatically from `kyutai/mimi` on first use.
## Training & evaluation
- **Method.** Built on the approach introduced in the [DualTurn paper](https://arxiv.org/abs/2603.08216) (frozen dual-channel Mimi
features β small causal turn-taking readout), specialized here for user end-of-turn endpointing.
- **Training data.** The two public dual-channel (humanβhuman) corpora used in the DualTurn paper β
**Switchboard** and **OtoSpeech** β plus a large **private dual-channel dataset** of real AI-agent phone
calls (the production domain this model targets).
- **Evaluation.** Measured on our **internal private test set** (held-out real agent calls; the recall /
interval and latency numbers above) and, for an external neutral comparison, on **LiveKit's
[eot-bench](https://huggingface.co/datasets/livekit/eot-bench-data)** (reported in
[Performance](#performance) above; other systems' rows are LiveKit's published numbers).
## Files
- `model.safetensors` β the ~1.5M-param readout + input-standardization stats (PyTorch paths).
- `config.json` β with `auto_map` for `trust_remote_code`.
- `modeling_dualturn.py` β PyTorch model + both PyTorch modes: **offline** (`model(wav, sr)`) and **streaming** (`model.streaming()`).
- `model.onnx` β self-contained **offline** ONNX graph (Mimi encoder + readout).
- `stream_tick.onnx` β self-contained **streaming** ONNX graph (one 80 ms tick, state in/out).
- `stream_tick_240.onnx` β 240 ms streaming graph (3 frames/tick, ~1.8Γ cheaper/sec of audio, strictly causal).
- `onnx_streaming.py` β `DualTurnONNXStreamer` (80 ms) + `DualTurnONNXStreamer240` (240 ms) helpers (pure `onnxruntime` + `numpy`).
## Notes & limitations
- This is a **perception** model (it emits turn-taking *signals*, not agent actions) β pair it with a thin
decision policy.
- **`eot_probs` is the user's end-of-turn** (this is an endpointing model). `vad_probs` and `fvad_probs`
cover **both** channels (user, agent).
- Best in conversational voice-agent (telephone) audio, the domain it was trained on.
- **Deps by path:** PyTorch modes (1, 2) need `torch` + `torchaudio` + `transformers`; ONNX modes (3, 4)
need only `onnxruntime` + `numpy` (+ `librosa`/`soundfile` for loading, `huggingface_hub` for download).
## Author
[Shangeth Rajaa](https://github.com/shangeth) β Senior ML Research Scientist, Anyreach AI.
## Citation
```bibtex
@misc{rajaa2026dualturnlearningturntakingdualchannel,
title={DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining},
author={Shangeth Rajaa},
year={2026},
eprint={2603.08216},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2603.08216},
}
```
|