Image Feature Extraction
Birder
PyTorch

Model Card for naflex_i_vit_b16_ap_c1_siglip-v2-webli

A NaFlex ViT B/16 image encoder from the SigLIP v2 model by Tschannen et al., converted to the Birder format for image feature extraction. This version preserves the original model weights and architecture for downstream tasks.

See: https://huggingface.co/google/siglip2-base-patch16-naflex for further details.

Model Details

  • Model Type: Image classification and detection backbone

  • Model Stats:

    • Params (M): 92.9
    • Input image size: 256 x 256
  • Papers:

Model Usage

Image Embeddings

import birder
from birder.inference.classification import infer_image

# Option 1: manual setup (more control over preprocessing)
net, model_info = birder.load_pretrained_model("naflex_i_vit_b16_ap_c1_siglip-v2-webli", inference=True)

# Get the image size the model was trained on
size = birder.get_size_from_signature(model_info.signature)

# Create a NaFlex inference transform
patch_size = net.stem_stride
max_seq_len = (size[0] // patch_size) * (size[1] // patch_size)
transform = birder.naflex_transform(patch_size, max_seq_len, model_info.rgb_stats)

# Option 2: helper (quick start with NaFlex preprocessing)
net, model_info, transform = birder.load_pretrained_model_and_transform(
    "naflex_i_vit_b16_ap_c1_siglip-v2-webli",
    inference=True,
    naflex=True,
)

image = "path/to/image.jpeg"  # or a PIL image
out, embedding = infer_image(net, image, transform, return_embedding=True)
# embedding is a NumPy array with shape of (1, 768)

Detection Feature Map

from PIL import Image
import birder

net, model_info, transform = birder.load_pretrained_model_and_transform("naflex_i_vit_b16_ap_c1_siglip-v2-webli", inference=True)

image = Image.open("path/to/image.jpeg")
features = net.detection_features(transform(image).unsqueeze(0))
# features is a dict (stage name -> torch.Tensor)
print([(k, v.size()) for k, v in features.items()])
# Output example:
# [('stage1', torch.Size([1, 768, 16, 16]))]

Citation

@misc{dosovitskiy2021imageworth16x16words,
      title={An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale},
      author={Alexey Dosovitskiy and Lucas Beyer and Alexander Kolesnikov and Dirk Weissenborn and Xiaohua Zhai and Thomas Unterthiner and Mostafa Dehghani and Matthias Minderer and Georg Heigold and Sylvain Gelly and Jakob Uszkoreit and Neil Houlsby},
      year={2021},
      eprint={2010.11929},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2010.11929},
}

@misc{tschannen2025siglip2multilingualvisionlanguage,
      title={SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features},
      author={Michael Tschannen and Alexey Gritsenko and Xiao Wang and Muhammad Ferjad Naeem and Ibrahim Alabdulmohsin and Nikhil Parthasarathy and Talfan Evans and Lucas Beyer and Ye Xia and Basil Mustafa and Olivier Hénaff and Jeremiah Harmsen and Andreas Steiner and Xiaohua Zhai},
      year={2025},
      eprint={2502.14786},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2502.14786},
}
Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for birder-project/naflex_i_vit_b16_ap_c1_siglip-v2-webli

Finetuned
(3)
this model

Papers for birder-project/naflex_i_vit_b16_ap_c1_siglip-v2-webli