Perch v2 (PyTorch)
An unofficial PyTorch port of Perch v2 (Perch 2.0), Google DeepMind's bioacoustics model for bird and wildlife sounds. The weights are a direct conversion of Google's official TensorFlow release (perch_v2, version 2, Apache 2.0), and the PyTorch model reproduces the TensorFlow outputs to float round-off.
Code, examples and documentation: https://github.com/matijama/perch-v2-pytorch
Usage
pip install git+https://github.com/matijama/perch-v2-pytorch
from perchv2_pytorch import PerchV2, load_audio, frame_audio
model = PerchV2.from_pretrained() # downloads these weights once, then cached
windows = frame_audio(load_audio("recording.wav")) # any sample rate -> 5 s windows at 32 kHz
out = model(windows)
out["embedding"] # [windows, 1536]
out["label"] # [windows, 14795] logits; class names in model.labels
out["spatial_embedding"] # [windows, 16, 4, 1536]
out["spectrogram"] # [windows, 500, 128]
model.predict(windows, top_k=3)[0]
For embeddings only, PerchV2.from_pretrained(embeddings_only=True) downloads just the 41 MB encoder. Half precision on a GPU needs no separate weights:
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.float16):
out = model.cuda()(windows.cuda())
The GitHub repository also has a fine-tuning example (train Perch's own head for your classes, a new head, or the whole model).
Files
| file | content |
|---|---|
model.safetensors |
full model: EfficientNet-B3 encoder + 14,795-class prototype classifier (388 MB, fp32) |
encoder.safetensors |
encoder only, for embeddings (41 MB) |
labels.csv |
index, label, ebird_code for the 14,795 output classes |
Citation
@article{vanmerrienboer2025perch2,
title = {Perch 2.0: The Bittern Lesson for Bioacoustics},
author = {van Merri{\"e}nboer, Bart and Dumoulin, Vincent and Hamer, Jenny and Harrell, Lauren and Burns, Andrea and Denton, Tom},
journal = {arXiv preprint arXiv:2508.04665},
year = {2025}
}
Licence
Apache License 2.0, the same as the original model. The weights are a format conversion of Google's release: the values are unchanged, and only the storage format and tensor layout differ.