Instructions to use espnet/voxcelebs12_ecapa_mel with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ESPnet
How to use espnet/voxcelebs12_ecapa_mel with ESPnet:
unknown model type (must be text-to-speech or automatic-speech-recognition)
- Notebooks
- Google Colab
- Kaggle
Add a usage example to the model card
Browse files
README.md
CHANGED
|
@@ -10,6 +10,19 @@ license: cc-by-4.0
|
|
| 10 |
pipeline_tag: audio-classification
|
| 11 |
---
|
| 12 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
## ESPnet2 SPK model
|
| 14 |
|
| 15 |
### `espnet/voxcelebs12_ecapa_mel`
|
|
|
|
| 10 |
pipeline_tag: audio-classification
|
| 11 |
---
|
| 12 |
|
| 13 |
+
## Usage
|
| 14 |
+
|
| 15 |
+
```python
|
| 16 |
+
import librosa
|
| 17 |
+
from espnet2.bin.spk_inference import Speech2Embedding
|
| 18 |
+
|
| 19 |
+
speech2embedding = Speech2Embedding.from_pretrained(model_tag="espnet/voxcelebs12_ecapa_mel")
|
| 20 |
+
# the speaker models are trained on 16 kHz mono; librosa gives that from
|
| 21 |
+
# whatever the file holds
|
| 22 |
+
speech, rate = librosa.load("audio.wav", sr=16000, mono=True)
|
| 23 |
+
embedding = speech2embedding(speech) # (1, embedding_dim)
|
| 24 |
+
```
|
| 25 |
+
|
| 26 |
## ESPnet2 SPK model
|
| 27 |
|
| 28 |
### `espnet/voxcelebs12_ecapa_mel`
|