VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking of Highly Similar Objects

Weights and precomputed assets for VLA-ReID (Video-Level Association re-ID): set-relative appearance re-scoring for multi-object tracking of highly similar objects, evaluated on the BEE24 benchmark. Frame-Common Appearance Estimation (FCAE) + Common Appearance Suppression (CAS) plug into TrackTrack by replacing only its appearance term.

BEE24 test (official 5-sequence split, shared YOLOX detections):

Method HOTA MOTA IDF1 AssA IDs
TrackTrack (baseline) 48.05 65.69 64.12 45.91 606
VLA-ReID (ours) 49.22 67.02 64.93 46.64 437

See the GitHub repository for installation and step-by-step reproduction instructions; the file-by-file manifest (destination paths + SHA-256) is in WEIGHTS.md.

File Bytes Role
yolox_x_bee24.pth.tar 792,835,731 YOLOX-X bee detector
bee24_test_0.80.pickle / bee24_test_0.95.pickle 6,162,932 / 8,818,700 shared detections (dual NMS)
sam3_bee24_finetuned.pt 3,371,888,621 SAM3 segmentation checkpoint fine-tuned on bee foreground (base model by Meta AI; redistributed under the SAM license terms)
bee_direction_resnet18_triclass.joblib 44,790,820 bee orientation classifier
reid_stage1_agw_s50.pth 331,487,317 stage-1 AGW-ResNeSt50 encoder (baseline)
reid_stage1_agw_s50_ep60.pth 331,487,317 stage-1 epoch-60 checkpoint (stage-2 init)
vla_reid_encoder.pth 102,195,690 stage-2 encoder export (feature extraction)
vla_reid_cas.pth 135,772,354 stage-2 FCAE+CAS checkpoint (tracker injection)
aflink.pth 4,348,705 AFLink post-linking weight

Code license: MIT (see the GitHub repo's NOTICE for vendored-component licenses). sam3_bee24_finetuned.pt is a derivative of Meta AI's SAM3 checkpoint and follows the SAM license.

If you use the BEE24 benchmark, please cite:

@ARTICLE{10851814,
  author={Cao, Xiaoyan and Zheng, Yiyao and Yao, Yao and Qin, Huapeng and Cao, Xiaoyu and Guo, Shihui},
  journal={IEEE Transactions on Image Processing},
  title={TOPIC: A Parallel Association Paradigm for Multi-Object Tracking Under Complex Motions and Diverse Scenes},
  year={2025},
  volume={34},
  pages={743-758},
  doi={10.1109/TIP.2025.3526066}}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support