File size: 5,843 Bytes
840d05a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
---
license: apache-2.0
language:
- zh
- yue
- en
library_name: onnx-asr
pipeline_tag: automatic-speech-recognition
tags:
- automatic-speech-recognition
- onnx
- onnx-asr
- paraformer
- funasr
- cantonese
base_model:
- csukuangfj/sherpa-onnx-paraformer-trilingual-zh-cantonese-en
---

# Paraformer trilingual ZH / Cantonese / EN ONNX

A SeACo-Paraformer trained on Mandarin, Cantonese (Yue) and English for onnx-asr.
This is the strongest available Cantonese model in this collection: it is character
identical to native FunASR on every Cantonese and Mandarin test clip, fp32 and int8
alike, and it replaces `OpenVoiceOS/whisper-small` as the Cantonese default, since
that model is a general multilingual Whisper checkpoint with no Cantonese-specific
training.

## What this is

Paraformer is the Alibaba [FunASR](https://github.com/modelscope/FunASR) offline
non-autoregressive recognizer. A SAN-M encoder reads the audio, a CIF predictor decides
how many tokens the utterance has, and a single pass decoder emits all of them at once.
There is no decoding loop, so one forward pass gives the transcript.

This checkpoint is a SeACo-Paraformer, the hotword-biasing variant of Paraformer. The
graph exposes two extra outputs used by the bias pathway (`decoder_hidden`,
`pre_acoustic_embeds`), but the `paraformer` model type in onnx-asr only requests
`logits` and `token_num` by name, so it loads and runs unmodified, with no hotword
biasing applied. Its `speech` / `speech_lengths` inputs and `logits` / `token_num`
outputs match the plain Paraformer contract exactly.

The ONNX graphs here are copied byte for byte from the
[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) export by
[csukuangfj](https://huggingface.co/csukuangfj/sherpa-onnx-paraformer-trilingual-zh-cantonese-en).
Only the side files changed: `tokens.txt` became `vocab.txt`, and `config.json` carries
the FunASR frontend statistics from `am.mvn`.

## Usage

The `paraformer` model type is on the `feat/paraformer` branch of the TigreGotico
[onnx-asr](https://github.com/TigreGotico/onnx-asr) fork.

```sh
pip install "onnx-asr[cpu,hub] @ git+https://github.com/TigreGotico/onnx-asr@feat/paraformer"
```

```py
import onnx_asr

model = onnx_asr.load_model("OpenVoiceOS/paraformer-trilingual-zh-yue-en-onnx")
print(model.recognize("audio.wav"))
```

## Graph contract

| Item | Value |
|---|---|
| Input | `speech`, float32, `[batch, num_frames, 560]` |
| Input | `speech_lengths`, int32, `[batch]` |
| Output | `logits`, float32, `[batch, num_tokens, 11666]` |
| Output | `token_num`, int32, `[batch]`, the CIF token count |
| Output (unused) | `decoder_hidden`, `pre_acoustic_embeds` — SeACo hotword pathway, not read by onnx-asr |

The 560 dim input is the FunASR frontend: an 80 dim kaldi fbank of a waveform scaled to
the int16 range, then a low frame rate stack of 7 frames with a hop of 6, then the
`am.mvn` mean variance statistics. onnx-asr computes the fbank with its `wespeaker`
preprocessor and applies the LFR stack and the CMVN in the runtime. Decoding is one
argmax per logits row, stopping at `</s>` and never reading past `token_num`.

## Streaming

A streaming Paraformer also exists upstream, using a different graph with encoder and
decoder states, and needs a streaming runtime that onnx-asr does not have yet (upstream
issue #21). Only the offline model is mirrored here.

## Files

`model.onnx` (871 MB) and `model_int8.onnx` (245 MB). Both come straight from the
sherpa-onnx repository; nothing was re-quantized here.

## Parity with native FunASR

2 Mandarin, 2 English and 2 Cantonese clips (FLEURS `cmn_hans_cn`, `en_us`,
`yue_hant_hk`), native FunASR 1.4.0 on the exact ModelScope checkpoint
(`dengcunqin/speech_seaco_paraformer_large_asr_nat-zh-cantonese-en-16k-common-vocab11666-pytorch`)
with `frontend.dither = 0`.

| Clip | fp32 | int8 |
| --- | --- | --- |
| zh_1 | identical | identical |
| zh_2 | identical | identical |
| en_1 | identical | identical |
| en_2 | identical | a few words differ (deliberately out-of-domain foreign-cuisine clip) |
| yue_1 | identical | identical |
| yue_2 | identical | identical |

6 of 6 fp32 comparisons and 5 of 6 int8 comparisons are character identical to native
FunASR. The one int8 difference is on the same deliberately hard, out-of-domain English
clip that shows quantization noise in the other Paraformer mirrors in this collection.
Cantonese, the language this checkpoint adds to the collection, is perfect in both
precisions.

## Speed

AMD Ryzen 5 7600, 6 cores, `OMP_NUM_THREADS=6`, CPU provider, batch 1, warm cache, best
of 3 passes over the 6 test clips.

| Precision | RTF |
| --- | --- |
| fp32 | 0.0196 |
| int8 | 0.0148 |

## License

Apache License 2.0. The [ModelScope model card](https://www.modelscope.cn/models/dengcunqin/speech_seaco_paraformer_large_asr_nat-zh-cantonese-en-16k-common-vocab11666-pytorch)
states `License: Apache License 2.0`. The sherpa-onnx mirror carries no license file of
its own, which is why the ModelScope source was read directly rather than the mirror.

* Model: Alibaba DAMO Academy / FunASR, `dengcunqin/speech_seaco_paraformer_large_asr_nat-zh-cantonese-en-16k-common-vocab11666-pytorch`, Apache-2.0.
* ONNX export: [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) by Fangjun Kuang (csukuangfj), Xiaomi.
* This repository only repackages those files for onnx-asr.

## Deviations

* The LFR stack and the CMVN are in the runtime class, not in the graph, matching every
  other Paraformer mirror in this collection.
* No timestamps. Paraformer has no per-token frame index.
* No inverse text normalization and no homophone replacement.
* The streaming Paraformer is out of scope (see above).
* This checkpoint is SeACo-Paraformer (hotword biasing), used here without hotwords.
  The bias-pathway outputs are present in the graph but unread.